Context-dependent utility and robustness of pretrained single-cell foundation model representations across analytical tasks
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Single-cell foundation models (scFMs) have emerged as powerful representation learning approaches for single-cell transcriptomics. However, the utility and robustness of their pretrained representations across diverse analytical tasks and data conditions remain insufficiently characterized, particularly in zero-shot settings without task-specific fine-tuning. Here, we systematically analyze zero-shot performance of single-cell transcriptomic representations across 20 methods, 6 downstream tasks and 1,607 datasets comprising nearly 21.8 million cells. We evaluate model behavior along three complementary dimensions: utility on original datasets, robustness to controlled changes in dataset structure, and exploratory associations between dataset characteristics and performance variation. Our results show that scFM performance is strongly task dependent, with no single method consistently outperforming others across cell- and gene-level analyses. Notably, high utility on original datasets did not necessarily translate into robustness under structural perturbations, and several top-ranking methods were sensitive to changes in cell number, gene number, class composition, class imbalance, and batch complexity. Conventional statistical and task-specific methods remained competitive in several settings, while greater computational cost did not consistently correspond to better performance. Driver analyses further identified task-specific associations between performance and dataset characteristics, including cell-type complexity, train-test class overlap, batch number, and regulatory target-set size. Together, these findings show that the zero-shot utility and robustness of pretrained scFM representations depend jointly on analytical task and dataset structure. Our study provides a practical basis for context-aware representation selection and underscores the importance of evaluating structural robustness alongside utility when developing and applying scFMs.