Whole genome similarity provides a rapid, robust framework for classification of fungal taxa from the genus rank to intraspecies variants
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Rapid and accurate microbial identification is critical for interpreting biological data in basic research and when making applied decisions on how to effectively treat patients and control human, animal, and plant diseases. Advancements in high-throughput sequencing have the potential to expedite fungal species identification and thus fungal biological research; however, analyses of ever-increasing numbers of genomes also present computational challenges. In this work, we evaluated how whole-genome similarity can serve as the basis for accurate classification and identification across the Kingdom Fungi and the Phylum Oomycota. Results from the computationally efficient k-mer–based tool sourmash are compared with those from more computationally demanding BLAST-based similarity method ANIb, as well as with conventional phylogenomic approaches, including maximum-likelihood concatenated ortholog trees and SNP-based methods. We observed that sourmash delivers orders-of-magnitude gains in speed and memory efficiency while maintaining strong concordance with phylogenomic methods. We found that a k-mer size of 21 is robust for genus and species determination, and that larger k-mers are well-suited for identification below the species rank. These results demonstrate that k-mer-based whole-genome similarity provides a scalable framework for fungal classification, enabling rapid analysis, lowering computational and bioinformatic barriers, and supporting the development of efficient identification pipelines.