Towards Principled Evaluation of Single-Cell Perturbation Prediction Models

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Single-cell perturbation experiments measure how interventions alter cellular phenotypes. However, the number of possible perturbations and biological contexts far exceeds what can be tested experimentally. Motivated by this constraint, predictive models aim to extrapolate cellular responses to unseen conditions. Despite substantial efforts in model development, benchmark studies have reached inconsistent conclusions about the capabilities of current perturbation-response models. A major challenge is that evaluation protocols vary widely across studies, making results difficult to compare. Furthermore, the lack of consensus on evaluation hampers progress because it is unclear which predictive capabilities new models should prioritize. To help build consensus on evaluation principles, we develop a taxonomy that decomposes evaluation protocols into their representation, metric, score transformation, and reporting strategies. We characterize how these choices determine which aspects of prediction quality a benchmark measures and discuss criteria for selecting and assessing protocols in relation to specific benchmarking goals. We additionally provide scPertEval, a Python package with reference implementations of selected evaluation protocols, and use it to assess protocol behavior across seven publicly available single-cell perturbation datasets. By making evaluation choices and their underlying trade-offs explicit, we aim to stimulate a community discussion about developing more comparable and task-aligned evaluation protocols.

Graphical Abstract

Article activity feed