Performance of protein panels is inflated across many biomarker studies

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Data leakage is a prevalent yet underappreciated flaw in biomarker discovery studies. Through simulation and real-world proteomic data, we demonstrate that typical pipelines are broadly susceptible to this issue, producing inflated performance estimates, poor generalization, and excess false positives. We further introduce two tools to detect data leakage at the code and manuscript level, providing a practical path toward more rigorous and reproducible biomarker reporting.

Article activity feed