Language model–assisted label refinement for accurate sepsis detection from electronic health records

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Sepsis is a leading cause of hospital mortality, yet timely recognition is hampered by nonspecific presentations and label noise in code-based case definitions. We developed STRIDE, a machine-learning framework for sepsis detection across seven hospitals with a scalable approach to label quality. We refined a pragmatic operational definition using a large language model applied to discharge summaries, with an independent physician-adjudicated cohort as the gold standard. We compared 8-, 24-, and 48-hour observation windows and benchmarked against SOFA, SIRS, and Epic, assessing calibration and discrimination. Among 356,610 encounters, the 8-hour model achieved an AUC of 0.960 in derivation and 0.878 in physician-adjudicated validation, matching or outperforming longer-window models. STRIDE outperformed SOFA and Epic on discrimination and showed favorable calibration by Brier score, while retaining strong discrimination among SIRS-positive non-septic encounters and reaching 78.2% specificity at 80% sensitivity in validation. These findings support accurate sepsis surveillance while limiting unnecessary alerts and requiring minimal prior history.

Article activity feed