Primer, Pipelines, Parameters: Issues in 16S rRNA Gene Sequencing

Isabel Abellan-Schneyder; Monica S Matchado; Sandra Reitmeier; Alina Sommer; Zeno Sewald; Jan Baumbach; Markus List; Klaus Neuhaus

doi:10.1128/mSphere.01202-20

Primer, Pipelines, Parameters: Issues in 16S rRNA Gene Sequencing

mSphere. 2021 Feb 24;6(1):e01202-20. doi: 10.1128/mSphere.01202-20.

Authors

Isabel Abellan-Schneyder¹, Monica S Matchado², Sandra Reitmeier¹, Alina Sommer¹, Zeno Sewald¹, Jan Baumbach^{2

3

4}, Markus List², Klaus Neuhaus⁵

Affiliations

¹ Core Facility Microbiome, ZIEL-Institute for Food & Health, Technische Universität München, Freising, Germany.
² Chair of Experimental Bioinformatics, TUM School of Life Sciences Weihenstephan, Technische Universität München, Freising, Germany.
³ Computational Biomedicine Lab, Department of Mathematics and Computer Science, University of Southern Denmark, Odense, Denmark.
⁴ Chair of Computational Systems Biology, University of Hamburg, Hamburg, Germany.
⁵ Core Facility Microbiome, ZIEL-Institute for Food & Health, Technische Universität München, Freising, Germany neuhaus@tum.de.

Abstract

Short-amplicon 16S rRNA gene sequencing is currently the method of choice for studies investigating microbiomes. However, comparative studies on differences in procedures are scarce. We sequenced human stool samples and mock communities with increasing complexity using a variety of commonly used protocols. Short amplicons targeting different variable regions (V-regions) or ranges thereof (V1-V2, V1-V3, V3-V4, V4, V4-V5, V6-V8, and V7-V9) were investigated for differences in the composition outcome due to primer choices. Next, the influence of clustering (operational taxonomic units [OTUs], zero-radius OTUs [zOTUs], and amplicon sequence variants [ASVs]), different databases (GreenGenes, the Ribosomal Database Project, Silva, the genomic-based 16S rRNA Database, and The All-Species Living Tree), and bioinformatic settings on taxonomic assignment were also investigated. We present a systematic comparison across all typically used V-regions using well-established primers. While it is known that the primer choice has a significant influence on the resulting microbial composition, we show that microbial profiles generated using different primer pairs need independent validation of performance. Further, comparing data sets across V-regions using different databases might be misleading due to differences in nomenclature (e.g., Enterorhabdus versus Adlercreutzia) and varying precisions in classification down to genus level. Overall, specific but important taxa are not picked up by certain primer pairs (e.g., Bacteroidetes is missed using primers 515F-944R) or due to the database used (e.g., Acetatifactor in GreenGenes and the genomic-based 16S rRNA Database). We found that appropriate truncation of amplicons is essential and different truncated-length combinations should be tested for each study. Finally, specific mock communities of sufficient and adequate complexity are highly recommended.IMPORTANCE In 16S rRNA gene sequencing, certain bacterial genera were found to be underrepresented or even missing in taxonomic profiles when using unsuitable primer combinations, outdated reference databases, or inadequate pipeline settings. Concerning the last, quality thresholds as well as bioinformatic settings (i.e., clustering approach, analysis pipeline, and specific adjustments such as truncation) are responsible for a number of observed differences between studies. Conclusions drawn by comparing one data set to another (e.g., between publications) appear to be problematic and require independent cross-validation using matching V-regions and uniform data processing. Therefore, we highlight the importance of a thought-out study design including sufficiently complex mock standards and appropriate V-region choice for the sample of interest. The use of processing pipelines and parameters must be tested beforehand.

Keywords: 16S rRNA gene sequencing; amplicon sequencing; bioinformatic settings; clustering; databases; microbiome; mock communities; variable regions.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Computational Biology
DNA Primers / genetics*
DNA, Bacterial / genetics*
Feces / microbiology
Gastrointestinal Microbiome / genetics*
Genetic Variation
High-Throughput Nucleotide Sequencing / methods
High-Throughput Nucleotide Sequencing / standards*
Humans
Phylogeny
RNA, Ribosomal, 16S / genetics*
Sequence Analysis, DNA

Substances

DNA Primers
DNA, Bacterial
RNA, Ribosomal, 16S