Dataset structure outweighs method choice in single-cell cell-type annotation
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Automated cell-type annotation is a prerequisite for nearly all single-cell RNA-sequencing (scRNA-seq) analysis, and the proliferation of annotation tools — spanning marker-based, similarity-based, classical machine-learning, deep-learning, semi-supervised, large language model (LLM), and transformer foundation-model families — has exceeded any existing guidance on how to select among them. Existing benchmarks utilize convenient, well-known datasets in which cell number, class imbalance, cell-type number, and differential-expression strength vary concurrently, so performance cannot be attributed to any dataset property. We benchmarked 63 tools from seven families using a Taguchi L9(3⁴) orthogonal array that varied these four properties independently and extended the comparison under more realistic conditions: real data, cross-platform transfer, public marker databases, LLM annotation, and fine-tuned foundation models. With standardized preprocessing and inputs, the leading cell-level families performed within 0.06 κ of one another, and fine-tuned foundation models averaged only slightly higher. Across families, sequencing platforms, and data types, accuracy was predicted near-linearly by how separable cell types were in a shared expression embedding space, measured as k-nearest-neighbor (kNN) purity (R² = 0.85–0.99). We found that dataset identity accounted for ≈84% of κ variance while tool identity accounted for ≈4%. Strong performers were distributed across families, for example, Seurat label transfer and singleCellNet among reference-based classifiers and mLLMCelltype among LLM annotators. When cell types cleanly separate, many annotation methods suffice; when they do not, algorithmic complexity struggles to compensate for noisy signal. Method selection should therefore prioritize available resources, such as curated references and compute availability, over sophisticated modeling.