GenomeCompendium: A database for the integrated analysis of repeats, assembly quality and functional content of complete prokaryotic genomes
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes ( https://genome-compendium.com/ ). Using NCBI RefSeq (∼47,000) and GenBank (∼13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classification and frequency analysis, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links ∼90 features across genomes, offers downloadable reports and -as unique features-pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.