Spectral universality in single-cell foundation models: a low-dimensional shared core carries the biology

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Dozens of foundation models are now trained on single-cell transcriptomes, each proposed as a general-purpose representation of the genome. We ask whether they agree with one another, and if so, where. Taking seven transcriptome-trained models spanning five architecture families and a 32-fold range of parameter counts, together with a protein language model as an outgroup, and restricting all of them to the 17,874 genes they share, we find that pairwise agreement is low: leading rank-256 subspaces overlap by 0.107 on average, where 1 is identity and 0.014 is chance. An empirical co-expression profile carrying no learned parameters overlaps each model more (0.135) than the models overlap each other. Generalised canonical correlation analysis nevertheless recovers a small consensus subspace: the leading shared direction lies 91% inside every model on average, and 43 directions lie at least half inside all of them, against a gene-permutation null of 0.22. This core is robust—dropping any model, or the entire three-member Geneformer family, leaves the leading shared fraction between 0.900 and 0.948. We then locate the biology. On five retrieval tasks built from Gene Ontology, Reactome, hu.MAP protein complexes and STRING, the component of each model orthogonal to the core is at or near chance (mean AUROC 0.501–0.551), while restricting a model to the core raises its mean AUROC for all eight models, significantly on all five tasks for seven of them. Eight consensus directions, belonging to no model, significantly outperform every full-width embedding on all five tasks (40 of 40 paired DeLong comparisons, minimum |z| = 7.6). The core survives removing a transcript-abundance axis and a co-expression subspace, both individually and jointly, and resolves into interpretable programs: a housekeeping-versus-untranscribed axis, then adhesion and extracellular matrix, cytokine signalling, immune response, and translational machinery. The practical consequence is that the majority of every model’s representational capacity is private and biologically inert, and the object worth extracting from these models is not any one model’s geometry but the intersection of several.

Article activity feed