A Sampling-Based Computational Frame-work for Evaluating Enterotype Assignment Stability in Clinical Microbiome Studies
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Enterotyping classifies individuals into discrete gut microbiome community types and is increasingly used for clinical stratification, yet whether individual assignments are reliable remains untested. This study introduces two resampling procedures to evaluate the stability of enterotype assignments derived from Dirichlet Multinomial Mixture (DMM) models: one perturbs the clinical composition of the cohort (ERP), the other its size (ERPM). Applied to an European cohort with shotgun metagenomics data (MetaCardis; N=2,022) over 500 iterations, the framework reveals that a substantial fraction of individuals change enterotype depending on the composition of the cohort. These shifts are not random: they follow reciprocal trade-offs between Bact1 and Bact2, and between Prevotella and Ruminococcus. The dysbiosis-associated Bact2 enterotype is the most sensitive, losing prevalence and confidence when diseased individuals are added to healthy cohorts, and gaining prevalence and confidence in the reverse scenario. Importantly, individuals fall into two groups: core individuals, whose assignments are stable regardless of context, and variable individuals, whose assignments depend on cohort composition. For variable individuals, collapsing the four enterotypes into two groups along the dominant trade-off axes restores reliable classification. Stable four-enterotype assignment requires large cohorts (>1,400 samples), with Ruminococcus consistently being the least robust. The framework provides practical tools for identifying which patients can be reliably enterotyped and which require alternative stratification.