Robust and Interpretable Metagenomic Modeling Through Structure-Aware Multi-View Learning and Attribution-Guided Biological Insight
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Integrative modeling of metagenomic and clinical data can advance the study of host phenotypes, but remains challenged by cross-view heterogeneity, uncertain generalizability, and poor interpretability. We developed SAMECAT (Structure-Aware Metagenomics multi-viEw Contrastive AlignmenT), a structure-aware deep learning framework that integrates species-level shotgun metagenomic profiles with mixed-type clinical covariates through view-specific encoders, clustering-informed contrastive alignment, and adaptive representation fusion. Using two independent Louisiana Osteoporosis Study datasets generated through distinct sequencing and bioinformatics pipelines (development n = 1,990; external evaluation n = 481), we evaluated SAMECAT for bone mineral density prediction at four skeletal sites. SAMECAT consistently outperformed single-view models, naive concatenation, alternative deep learning integration approaches, and established machine learning baselines, with performance gains largely preserved in cross-pipeline external evaluation. To improve biological interpretability, we developed a stability-oriented interpretation workflow that aggregates individually low-magnitude and diffusely distributed feature attributions into structured modules, revealing reproducible site-dependent patterns, coherent functional themes, and representative hub taxa. SAMECAT thus provides a robust and interpretable framework for multi-view metagenomic modeling of microbiome-associated host phenotypes.