Edge effects in calling variants from targeted amplicon sequencing

Ravi Vijaya Satya; John DiCarlo

doi:10.1186/1471-2164-15-1073

Edge effects in calling variants from targeted amplicon sequencing

BMC Genomics. 2014 Dec 5;15(1):1073. doi: 10.1186/1471-2164-15-1073.

Authors

Ravi Vijaya Satya¹, John DiCarlo

Affiliation

¹ Research and Foundation Department, QIAGEN Sciences, Inc,, Frederick, MD, USA. ravi.vijayasatya@qiagen.com.

Abstract

Background: Analysis of targeted amplicon sequencing data presents some unique challenges in comparison to the analysis of random fragment sequencing data. Whereas reads from randomly fragmented DNA have arbitrary start positions, the reads from amplicon sequencing have fixed start positions that coincide with the amplicon boundaries. As a result, any variants near the amplicon boundaries can cause misalignments of multiple reads that can ultimately lead to false-positive or false-negative variant calls.

Results: We show that amplicon boundaries are variant calling blind spots where the variant calls are highly inaccurate. We propose that an effective strategy to avoid these blind spots is to incorporate the primer bases in obtaining read alignments and post-processing of the alignments, thereby effectively moving these blind spots into the primer binding regions (which are not used for variant calling). Targeted sequencing data analysis pipelines can provide better variant calling accuracy when primer bases are retained and sequenced.

Conclusions: Read bases beyond the variant site are necessary for analysis of amplicon sequencing data. Enzymatic primer digestion, if used in the target enrichment process, should leave at least a few primer bases to ensure that these bases are available during data analysis. The primer bases should only be removed immediately before the variant calling step to ensure that the variants can be called irrespective of where they occur within the amplicon insert region.

MeSH terms

Computational Biology / methods*
Computer Simulation
DNA Primers
High-Throughput Nucleotide Sequencing*
Polymerase Chain Reaction / methods
Reproducibility of Results
Sequence Analysis, DNA / methods*

Substances

DNA Primers