CytoGate-Bench: an LLM benchmark for cross-panel cell gating in cytometry

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

In cytometry, the workhorse single-cell technology of clinical immunology, every study defines its own antibody panel and cell-type vocabulary, so a classifier trained on one cannot annotate the next. Immunologists instead annotate by manual gating, splitting one parent population at a time on a two-marker plot, down an expert-defined hierarchy. We introduce CytoGate-Bench, a benchmark that reformulates this per-step procedure as a zero-shot, panel-agnostic task for large language models. It comprises 23,646 expert-annotated instances re-curated from 11 public flow- and mass-cytometry cohorts spanning eight marker panels. Across six open- and closed-weight backbones, the strongest formulation draws one rectangular gate per candidate and falls within the range of trained, panel-specialized baselines. It degrades less under distribution shift. Walking the hierarchy stepwise outperforms predicting every cell type at once. Ablations trace the signal to the data distribution shape and curated marker priors. However, adding vision or a self-verification loop systematically tightens gates.

THE BIGGER PICTURE

Immunology laboratories worldwide profile blood and tissue with cytometry, an instrument family that measures dozens of protein markers on millions of individual cells. Before any biology can be read out, every cell must be assigned an identity, a step still dominated by manual “gating,” in which an expert draws boundaries on a sequence of two-marker plots, following a documented, hierarchical protocol. Automating this step has remained difficult because every study measures a different marker panel and names a different set of cell types, so conventional machine-learning models must be retrained for each new study.

Large language models (LLMs) promise a different route, a single general-purpose model that reads the expert’s protocol and the data and makes each gating decision directly, with no study-specific training. This work contributes a public benchmark that tests precisely that ability across 11 human cohorts. The result is a statement of feasibility rather than superiority. Off-the-shelf models already score in the range of study-specific trained models and tolerate the day-to-day variation that degrades them. That capability matters most for new or small studies, for which no labeled training data exists. Walking the expert’s hierarchy one decision at a time also outperforms asking the model to name every cell type in a single pass, evidence that the structure of expert practice matters more than the scale of the question. These results come from a deliberately minimal setup, untuned models drawing simple rectangular gates, so we read them as a floor rather than a ceiling. Cytometry-aware training, richer gate geometries, and better-calibrated visual feedback are open avenues, and the benchmark gives that progress a fixed yardstick. Sustained progress would give laboratories analysts that keep pace with evolving marker panels without retraining, while leaving a decision trail an immunologist can audit.

Article activity feed