Perturbation response decomposition enables biologically aligned generalization to unseen perturbations and cellular contexts
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Predicting single-cell responses to genetic perturbations could reveal the vast combinatorial space of perturbations and cellular contexts that is infeasible to measure experimentally, yet current deep learning models generalize poorly and often fail to outperform simple baselines. Here we demonstrate that generalizability in perturbation prediction requires identifying and representing distinct components of cellular response rather than on increasing model complexity alone. We introduce a decomposition framework that explicitly separates transcriptional responses into global, perturbation-specific, cell-line-specific, and perturbation-by-cell-line interaction components. Applied to four CRISPR-interference Perturb-seq screens on multiple cell lines, our framework reveals that these components have distinct structures and information requirements. The global response component is low-dimensional, reflects recurrent proliferation and stress response programs, and can be inferred from control gene expression. In contrast, the perturbation and cell-line specific components are high-dimensional and cannot be recovered from control expression. We therefore develop response-component-aligned models that map biological priors, such as gene coessentiality, onto the geometry of observed transcriptional responses. Critically, this alignment enables simple linear or multilayer perceptron (MLP) based models to outperform state-of-the-art architectures across multiple generalization settings, including unseen cell lines and combinations of unseen perturbations. Together, our framework for response decomposition and alignment provides a principled basis for evaluating and designing perturbation-prediction models, showing that generalization depends primarily on matching biological information to the response components rather than on model complexity alone.