Systematic benchmark of ancient DNA read mapping

Adrien Oliva; Raymond Tobler; Alan Cooper; Bastien Llamas; Yassine Souilmi

doi:10.1093/bib/bbab076

Systematic benchmark of ancient DNA read mapping

Brief Bioinform. 2021 Sep 2;22(5):bbab076. doi: 10.1093/bib/bbab076.

Authors

Adrien Oliva¹, Raymond Tobler¹, Alan Cooper², Bastien Llamas^{1

3}, Yassine Souilmi^{1

3}

Affiliations

¹ Australian Centre for Ancient DNA, School of Biological Sciences, The University of Adelaide, South Australia, 5005, Australia.
² South Australian Museum, Adelaide, SA 5005, Australia.
³ The Environment Institute, The University of Adelaide, South Australia, 5005, Australia.

PMID: 33834210
DOI: 10.1093/bib/bbab076

Abstract

The current standard practice for assembling individual genomes involves mapping millions of short DNA sequences (also known as DNA 'reads') against a pre-constructed reference genome. Mapping vast amounts of short reads in a timely manner is a computationally challenging task that inevitably produces artefacts, including biases against alleles not found in the reference genome. This reference bias and other mapping artefacts are expected to be exacerbated in ancient DNA (aDNA) studies, which rely on the analysis of low quantities of damaged and very short DNA fragments (~30-80 bp). Nevertheless, the current gold-standard mapping strategies for aDNA studies have effectively remained unchanged for nearly a decade, during which time new software has emerged. In this study, we used simulated aDNA reads from three different human populations to benchmark the performance of 30 distinct mapping strategies implemented across four different read mapping software-BWA-aln, BWA-mem, NovoAlign and Bowtie2-and quantified the impact of reference bias in downstream population genetic analyses. We show that specific NovoAlign, BWA-aln and BWA-mem parameterizations achieve high mapping precision with low levels of reference bias, particularly after filtering out reads with low mapping qualities. However, unbiased NovoAlign results required the use of an IUPAC reference genome. While relevant only to aDNA projects where reference population data are available, the benefit of using an IUPAC reference demonstrates the value of incorporating population genetic information into the aDNA mapping process, echoing recent results based on graph genome representations.

Keywords: alignment; ancient DNA; benchmarking; reference bias.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Algorithms
Benchmarking / methods*
Computational Biology / methods*
DNA, Ancient / analysis*
DNA, Ancient / chemistry
Genome, Human / genetics*
High-Throughput Nucleotide Sequencing / methods
Humans
Polymorphism, Single Nucleotide
Reproducibility of Results
Sequence Alignment / methods*
Sequence Analysis, DNA / methods*
Software

Substances

DNA, Ancient