MIDFA: Scalable Bayesian Factor Analysis for Mixed and Incomplete Data

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Probabilistic latent variable models are a powerful tool for uncovering structure in high-dimensional datasets, particularly in biomedical applications. The increasing availability of large-scale epidemiological studies, such as the UK Biobank, poses important modelling challenges, including mixed data types, high dimensionality, and structured missingness. Existing approaches address some of these issues, but few provide a unified and scalable framework for handling them simultaneously. Here, we propose a scalable Bayesian factor analysis framework designed to address these challenges. Our method combines a semi-parametric Gaussian copula model with a continuous spike-and-slab prior to induce sparse and interpretable factor loadings. The number of latent dimensions is learned nonparametrically from the data using an Indian buffet process prior. For model fitting, we develop an expectation–maximisation algorithm that naturally accommodates missing data. We validate the proposed method through comprehensive simulation studies. In addition, we showcase the proposed model using the Novartis–Oxford Multiple Sclerosis dataset in two ways. First, we identify latent dimensions shared across MS clinical and neuroimaging variables, characterising disease structure while demonstrating the model’s ability to handle multiple data types and structured missingness. Second, we use the model for dimensionality reduction of structural MRI data, extracting features for downstream analysis that go beyond traditional whole-brain summary statistics. In those applications, our method identifies sparse latent structures and provides insights beyond those obtained from traditional approaches.

Article activity feed