Learning Minimal Gene Programs for Disease-Aligned Representations
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Identifying small, interpretable gene sets that robustly capture disease-associated variation in singlecell transcriptomic data remains a central challenge for biological interpretation and experimental followup. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding.
We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions.
Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-heldout evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data.
Data and Code Availability
All code used for data preprocessing, model training, evaluation, and feature importance analyses is available at: https://github.com/AdiVM/SLMC_single-cell . The singlecell data used in this study are publicly available human transcriptomic datasets generated by prior studies and accessible through the Gene Expression Omnibus (GEO). Analyses were performed using renal cell carcinoma single-cell RNA-seq data (GEO accession: GSE314072 ) and Alzheimer’s disease single-nucleus RNA-seq data from human cortex (GEO accession: GSE138852 ), with additional publicly available datasets used for cross-context diagnostic evaluation (GEO accessions: GSM8652069 , GSE308624 , and GSE227734 ). All datasets contain de-identified human samples and are available under standard publicuse terms via GEO.
Institutional Review Board (IRB)
This study analyzes de-identified, publicly available human transcriptomic data obtained from previously published studies. No new data were collected, and no identifiable private information was accessed. In accordance with institutional policy, this work was determined to constitute non-human subjects research and did not require additional IRB approval.