Mandatory use of mock communities highlighted by the descriptive comparison of Epi2Me 16S and EMU bioinformatic workflows for full-length 16S rRNA Nanopore sequencing

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Full-length 16S rRNA gene sequencing using Oxford Nanopore Technologies has emerged as a promising approach to improve species-level resolution in microbiota studies. However, the accuracy of taxonomic assignment remains highly dependent on the bioinformatics s used to process Nanopore long-read data. Therefore, the only way to ensure a good level of certainty in obtained results is to use positive controls in the form of mock communities in the experimental designs. In this study, we compared the performance of Epi2Me 16S (using Minimap2 or Kraken2) workflows provided by Oxford Nanopore Technologies and an EMU workflow for full-length 16S rRNA gene analysis. Using a commercial mock community sequenced across multiple Nanopore runs, taxonomic assignment accuracy and reproducibility was evaluated. Epi2Me-Kraken2 exhibited 18 % of incorrect genus-level assignments and failed to identify 3 species present in the mock community. While Epi2Me-Minimap2 achieved an excellent genus-level classification, reporting 9 % of sequences assigned to a genus not in the mock community, species-level assignments were inconsistent for several community members such as Listeria . In contrast, EMU provided accurate and consistent species-level taxonomic profiles, with all species correctly identified while keeping the number of genus absent from the mock community at 1.2%.

Importance

These results highlight that Epi2Me integrated workflows are not the best option for specie-level taxonomic assignation. More importantly, this paper underscores the importance of routine inclusion of positive controls for microbiota studies, in the form of mock communities, as a critical safeguard for accurate data interpretation. Without the use of a mock community, a paper published would be at risk of reporting wrong observations and inaccurate conclusions.

Article activity feed