A hyperspherical deep Bayesian model for interpretable clustering and relationship prediction in microbiome multi-omics integration

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

The microbiome plays a significant role in the development and progression of many diseases, yet extracting interpretable insights from multi-omics data remains challenging. Existing approaches face a recurring practical trade-off: deep learning methods achieve high predictive performance but lack uncertainty quantification, whereas probabilistic methods provide interpretable results but require data-type-specific likelihood functions that limit generalization across diverse omics modalities. Here, we introduce DBayesCM (Deep Bayesian Clustering for Multi-omics), which combines deep learning modeling with Bayesian nonparametric methods. DBayesCM employs separate encoders to project microbiome and host omics data into a shared latent space, where an infinite mixture model with a Dirichlet process prior determines the number of clusters automatically while quantifying the uncertainty of each sample’s assignment. Spike-and-slab priors identify discriminative features, and a Bayesian neural network estimates probabilistic co-occurrence between microbial species and host omics features. To isolate the effect of latent geometry, we evaluate two variants that are identical except for their latent space: DBayesCM-vMF constrains the latent to the unit hypersphere and applies a von Mises-Fisher mixture, while DBayesCM-GMM uses a Euclidean latent space and a Gaussian mixture. On simulated data, the hyperspherical variant recovered the correct number of clusters, whereas the Euclidean variant over-segmented, demonstrating that the latent geometry affects cluster recovery. Applied to colon, breast, and kidney cancer cohorts spanning metagenomics, host metabolomics, RNA-seq, and miRNA data, and to an obstructive sleep apnea model, DBayesCM ranked consistently among the existing methods while uniquely combining data-driven cluster-number determination, sample-level uncertainty, and interpretable feature selection within a single framework. DBayesCM reveals conditional probabilistic co-occurrence between core microbial species and host omics features, enabling uncertainty-aware exploration of microbiome-host relationships across diverse diseases.

Article activity feed