Class-Support Mismatch Dominates Protocol-Driven Optimism in Spatial Transcriptomics, with a Larger Residual Penalty for Graph Neural Networks

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Motivation: Spatial transcriptomics (ST) benchmarks are routinely reported to be “inflated” by randomly interleaved cross-validation, with the inflation attributed to spatial leakage. That attribution has never been tested. The random-versus-block performance gap simultaneously changes the training-set size, the set of classes available in training and test, the amount of message passing that crosses the train/test boundary, and the scope over which preprocessing is fit. A single gap statistic cannot separate them. Results: We build a graph in which train, validation and test induced subgraphs are mutually disconnected—zero cross-role edges and infinite minimum train–test hop distance in every fold, verified by two independent implementations—and then walk a ladder of eleven protocols that change one design axis at a time on 10x Visium human breast cancer (3,798 spots, 11 Leiden domains; 6 methods × 5 folds × 5 seeds). Folds 1 and 3 are the same spatial partition, so all inference uses n = 4 independent spatial blocks with hierarchical bootstrap intervals. The closed- set random-to-block gap of +0.447 macro-F1 [95% CI +0.332, +0.562] decomposes additively into class-support mismatch +0.276 [+0.121, +0.429], residual spatial extrapolation +0.098 [+0.065, +0.135], preprocessing scope +0.051 [+0.011, +0.096] and training-set size +0.022 [+0.014, +0.029]. Class-support mismatch, not leakage, is the largest single component. After matching size, class support, graph masking rule and preprocessing scope, the genuine spatial- extrapolation residual is +0.214 [+0.160, +0.273] for graph neural networks but +0.042 [−0.039, +0.143] for non-graph classifiers, a family gap of +0.172 [+0.054, +0.274]. Cross-role message passing is worth +0.050 (GCN) and +0.016 (GAT) macro-F1 and is exactly 0.000 for every non-graph method by construction. Recomputing persistent-homology features within each fold removes the previously reported +0.0154 topology margin entirely: the leakage channel alone accounts for +0.0141 of it, and the clean estimate is −0.0144 (SD 0.0269, exact sign-permutation p = 0.50 against a floor of 0.125), i.e. no difference detected in either direction. On 12 expert- annotated DLPFC sections (47,280 spots, 3 donors) the mixed-model protocol effect is −0.5488 (SE 0.0089) and the method ranking reverses: XGBoost-PCA+XY beats SVM-PCA in 12/12 sections under random CV and loses under block CV (2/12), leave-one-section-out (1/24) and leave-one-donor-out (1/6). Two published methods behave the same way (STAGATE +100.7%, GraphST +62.4%). Conclusion: Protocol-driven optimism is not one quantity. In the deeply audited breast-cancer case study the largest component is a class-support artifact of spatial blocking; the residual spatial-extrapolation penalty that survives full matching is substantially larger for graph neural networks and is statistically resolvable only for them, while for non-graph classifiers it is small and unresolved. Benchmarks should report the decomposition, not a single inflation figure, and should match class support before attributing a gap to space. Availability: Code, fixed split definitions, pruned graphs, machine-readable result tables and scripts for reproducing all analyses and figures are available at https://github.com/ChimdiWalter/Topospatial,ChimdiWalter/Topospatial. The exact release corresponding to this manuscript is deposited in the Zenodo archive.

Article activity feed