PathwayBench: a multi-criterion benchmark of pseudobulk pathway activity scoring methods reveals rank-window competition as a mechanism of biological signal loss in single-cell RNA-seq

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background Pathway activity scoring is a foundational step in single-cell RNA-seq (scRNA-seq) analysis, yet method choice is rarely guided by systematic, multi-criterion evidence in the pseudobulk case-control regime that now dominates applied single-cell disease studies. Prior single-cell pathway-scoring benchmarks have focused on perturbation ground truth at cell-level resolution, leaving the donor-level pseudobulk setting — and the question of why methods disagree — largely unaddressed. Results We present PathwayBench, a benchmark comparing five widely used pathway scoring methods (ssGSEA, GSVA, z-score, AUCell, UCell) across eight pseudobulked scRNA-seq case-control datasets spanning five tissues (brain, heart, kidney, lung, blood) and 682 donors. Methods are evaluated against five criteria covering biological relevance (direction accuracy, AUROC, effect size) and four robustness axes (aggregation, outlier, normalization, sample-size stability). We identify rank-window competition, a mechanism by which rank-based methods (AUCell, UCell) lose biological signal — and in extreme cases invert it — when non-pathway competitor genes saturate the top-rank scoring window. Controlled simulations demonstrate sign inversion in 40% of replicates under high competitor burden, and real-data confirmation comes from extracellular matrix (ECM) remodeling in chronic kidney disease, where rank-based methods produce wrong-direction effects. Discretizing per-criterion performance as Good/Intermediate/Poor reveals that no single method satisfies all five criteria, with GSVA and UCell each meeting four along different axes. Conclusions Rank-window competition provides a mechanistic explanation for systematic divergence between magnitude-aware and rank-based pathway scoring methods on scRNA-seq data and should guide method selection, particularly for fibrotic, inflammatory, or otherwise broadly remodeled transcriptomes. PathwayBench is delivered as a versioned, extensible benchmark with per-dataset scores, an interactive advisor application, and complete reproducibility infrastructure to support evidence-based method selection and community extension.

Article activity feed