An in-depth update on the benchmarks for 16S amplicon sequencing

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Amplicon-based techniques provide a rapid and cost-effective approach for profiling microbial communities. However, the observed microbial diversity is influenced by a wide range of factors, encompassing pre-analytical steps such as the choice of primers and target regions, as well as the bioinformatic pipeline, including the selection of tools, reference databases, and parameter settings. Several benchmarks are already available in the literature, but the updates to important tools and databases, namely LotuS3, the Ribosomal Database Project and GreenGenes2, prompted our investigation. In this study, we conducted a comprehensive benchmark of the main bioinformatic tools and databases. Using seven regions for three publicly available mock communities of increasing complexity, we tested 38 possible combinations of sequence resolution algorithms (DADA2 stand-alone, LotuS3 (DADA2/UPARSE)), taxonomic classifiers and search tools (Kraken2, DECIPHER, RDP, MMseqs2, Lambda, and Metaxa2), and databases (SILVA, GreenGenes2, RDP, RefSeq, and Metaxa2). The region V1-V3, coupled with DADA2+MMseqs2+SILVA, DADA2+Metaxa2, or LotuS3 (DADA2)+RDP yielded the highest-quality estimates of the true diversity according to the metrics. We also demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.

Article activity feed