Extracting abundance information from DNA-based data

Mingjie Luo; Yinqiu Ji; David Warton; Douglas W Yu

doi:10.1111/1755-0998.13703

Extracting abundance information from DNA-based data

Mol Ecol Resour. 2023 Jan;23(1):174-189. doi: 10.1111/1755-0998.13703. Epub 2022 Aug 30.

Authors

Mingjie Luo^{1

2}, Yinqiu Ji¹, David Warton^{3

4}, Douglas W Yu^{1

5

6}

Affiliations

¹ State Key Laboratory of Genetic Resources and Evolution and Yunnan Key Laboratory of Biodiversity and Ecological Security of Gaoligong Mountain, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming, Yunnan, China.
² Kunming College of Life Sciences, University of Chinese Academy of Sciences, Kunming, Yunnan, China.
³ School of Mathematics and Statistics, UNSW Sydney, Sydney, New South Wales, Australia.
⁴ Evolution and Ecology Research Centre, UNSW Sydney, Sydney, New South Wales, Australia.
⁵ Center for Excellence in Animal Evolution and Genetics, Chinese Academy of Sciences, Kunming, Yunnan, China.
⁶ School of Biological Sciences, University of East Anglia, Norwich Research Park, Norwich, Norfolk, UK.

Abstract

The accurate extraction of species-abundance information from DNA-based data (metabarcoding, metagenomics) could contribute usefully to diet analysis and food-web reconstruction, the inference of species interactions, the modelling of population dynamics and species distributions, the biomonitoring of environmental state and change, and the inference of false positives and negatives. However, multiple sources of bias and noise in sampling and processing combine to inject error into DNA-based data sets. To understand how to extract abundance information, it is useful to distinguish two concepts. (i) Within-sample across-species quantification describes relative species abundances in one sample. (ii) Across-sample within-species quantification describes how the abundance of each individual species varies from sample to sample, such as over a time series, an environmental gradient or different experimental treatments. First, we review the literature on methods to recover across-species abundance information (by removing what we call "species pipeline biases") and within-species abundance information (by removing what we call "pipeline noise"). We argue that many ecological questions can be answered with just within-species quantification, and we therefore demonstrate how to use a "DNA spike-in" to correct for pipeline noise and recover within-species abundance information. We also introduce a model-based estimator that can be used on data sets without a physical spike-in to approximate and correct for pipeline noise.

Keywords: Arthropoda; DNA barcoding; Insecta; biomonitoring; community composition; environmental DNA; internal standard; polymerase chain reaction; taxonomic bias.

Publication types

Review

MeSH terms

Biodiversity
DNA / genetics
DNA Barcoding, Taxonomic* / methods
Metagenomics* / methods

Substances

DNA

Abstract

Publication types

MeSH terms

Substances

Grants and funding