MGM2 as a Unified Foundation Model for Microbiome World Exploration
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Microbiomes are information-rich biological systems, yet most computational analyses still reduce communities to cohort-specific abundance tables. Here we introduce MGM2, a multimodal foundation model pretrained on 1,821,291 MicrobeAtlas samples and 225,067 OTUs clustered at 99% sequence similarity. MGM2 couples NTv3-derived microbial sequence embeddings with abundance conditioning and community-semantic alignment to learn transferable sample– and token-level representations. Frozen MGM2 representations outperformed DeepPhylo by 0.06–0.21 macro-AUROC across five temporally held-out MGnify hierarchy levels, with the largest gains for rare and fine-grained labels. In fecal microbiota transplantation, MGM2-XLarge achieved a response ROC AUC of 0.79 and reduced post-transplant Bray-Curtis distance by 15% relative to the recipient baseline. The same representation supported ASV-level trend forecasting across 24 wastewater treatment plants. Sparse autoencoder analysis resolved MGM2-XLarge token states into a 4,096-feature dictionary spanning taxonomic identity, abundance state, ecological context and technical variation. MGM2 therefore provides a sequence-aware and interpretable representation layer for microbiome classification, paired-community prediction, forecasting and feature discovery.
Highlights
MGM2 integrates sequence, abundance and community semantics through pretraining on 1.82 million microbiome samples.
Frozen MGM2 improved macro-AUROC over DeepPhylo by 0.06–0.21 across five temporally held-out MGnify levels.
MGM2-XLarge reached a response ROC AUC of 0.79 and reduced post-FMT Bray-Curtis distance by 15%.
A 4,096-feature sparse autoencoder atlas resolves taxonomic, abundance, ecological and technical signals.