Beyond P-values: A Multi-Metric Framework for Robust Feature Selection and Predictive Modeling

Raelynn Chen
Attri Ghosh
Jie Hu
Yong Chen
Jason H. Moore
Ruowang Li

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does not guarantee predictive utility, and vice versa. Yet few methods unify inferential and predictive evidence within a single selection framework. We introduce MIXER (Multi-metric Integration for eXplanatory and prEdictive Ranking), a domain-agnostic approach that integrates multiple selection metrics into one consensus model via adaptive weighting that quantifies each criterion’s contribution. Through simulation studies, we demonstrate that different selection metrics identified markedly different feature sets whose overlaps depended on the underlying feature distributions and signal strength. Applied to Alzhemier’s disease in UK Biobank, MIXER outperformed every individual criterion, including statistical significance, and generalized to an external disease-specific cohort, Alzheimer’s Disease Sequencing Project, yielding higher discrimination and stronger risk stratification. The MIXER framwork is also modular and readily extends to other selection criteria and data modalities, providing a practical route to more accurate, interpretable, and transportable predictive models.

Version published to 10.1101/2025.10.05.680380 on bioRxiv
Oct 6, 2025

Multi-Perspective Machine Learning MPML: A High-Performance and Interpretable Ensemble Method for Heart Disease Prediction

This article has 6 authors:
1. Sean Miller
2. Keaton Logan
3. Ricardo Anderson
4. Patricia Cowell
5. Curtis Busby-Earle
6. Lisa-Dionne Morris
This article has no evaluationsLatest version Aug 21, 2025
Coxmos: Interpretable survival models for high-dimensional and multi-omic data

This article has 3 authors:
1. Pedro Salguero
2. Anabel Buendía-Galera
3. Sonia Tarazona
This article has no evaluationsLatest version Oct 3, 2025
Reliable Bayesian Network Structure Learning in Biomedical Applications: Model Uncertainty Criterion and Its Operating Characteristics

This article has 2 authors:
1. Grigoriy Gogoshin
2. Andrei S Rodin
This article has no evaluationsLatest version Oct 14, 2025

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Multi-Perspective Machine Learning MPML: A High-Performance and Interpretable Ensemble Method for Heart Disease Prediction

Coxmos: Interpretable survival models for high-dimensional and multi-omic data

Reliable Bayesian Network Structure Learning in Biomedical Applications: Model Uncertainty Criterion and Its Operating Characteristics