Where in the spectrum does biology live? A confound audit and a compaction law for gene embeddings of single-cell foundation models
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
In language models, the leading directions of the embedding spectrum encode how often a token appears rather than what it means. Single-cell foundation models (scFMs) tokenise non-zero expression counts, so they have an exact analogue of token frequency: how abundantly a gene is expressed and in how many cells it is detected. This paper asks two questions of the gene-embedding tables of eight released models — seven single-cell transformers and the frozen ESM2 protein language model that Universal Cell Embeddings uses as input — over a common universe of 17,846 genes. (A) A confound audit. Regressing every spectral axis on a spline basis in abundance and detection breadth shows that the leading axis of six of the seven count-trained models is dominated by abundance (R2 = 0.51–0.76 against a permutation null of 7 × 10−4 ), while the protein language model that never saw an expression value sits at R2 = 0.009. The loading is concentrated: beyond axis 64 it falls below 0.01. We then show that deleting the leading axes improves curated-relation retrieval — the “all-but-thetop” effect, replicated in gene embeddings — and that band-limited retrieval peaks in the spectral band [32, 64), not at the top. Under abundance-matched negatives, roughly a third of the apparent complex- and pathway-retrieval margin over chance disappears, a co-expression control is almost entirely abundance, and regulator–target retrieval is unaffected. A supervised 29-way compartment probe reaches 4.4× chance while an abundance-only classifier reaches exactly chance, so compartment structure is real — but no single axis carries it. (B) A compaction law. For each relation type we measure the bandwidth k∗ , the smallest number of leading axes recovering 95% of a model’s own ceiling. The hypothesis that k∗ is a property of the biology is rejected: a two-way analysis of variance attributes η^2 = 0.44 to the relation and η^2 = 0.34 to the model under the standard protocol, and under abundance matching the model term dominates (0.62 versus a non-significant 0.17). The model effect is not explained by embedding width, and a cluster bootstrap shows the across-model spread exceeds sampling noise for pathway and complex membership. What transfers better is the ordering: the ranking of relations by bandwidth agrees across architectures (Kendall’s W = 0.64, p = 0.004). Compaction curves follow a stretched exponential with exponent at or below one, and a rank-64 truncation retains 0.85–0.95 of the ceiling on average, with the worst individual cases near 0.74. Both confound controls raise k∗ , so bandwidths measured without them under-provision rank. The practical recommendations: report the abundance R2 of any axis claimed to encode a biological property, evaluate with abundance-matched negatives, and choose a truncation rank from the relation the surrogate must carry rather than from a single tuned number.