Aligning Reinforcement Learning with Clinical Practice for Safe Decision Support in Pediatric Sepsis

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Offline reinforcement learning (RL) has emerged as a promising framework for clinical decision support in sepsis, yet most existing studies focus exclusively on adult populations, leaving pediatric care largely unexplored despite important physiological and treatment differences. In this work, we develop offline RL policies for pediatric sepsis management in the Pediatric Intensive Care Unit (PICU) using a retrospective cohort of 2,229 episodes from Great Ormond Street Hospital (GOSH), formalized as finite-horizon Markov Decision Process (MDP) with joint intravenous fluid and vasopressor actions. To better capture pediatric organ dysfunction dynamics, we incorporate Phoenix-8, a recently proposed pediatric sepsis severity score, as an intermediate reward shaping signal in addition to terminal 90-day mortality. We systematically vary the time-step size (4, 8, and 12 hours) and reward structure (terminal 90-day mortality, with and without Phoenix-8–based intermediate shaping), and compare Double Deep Q-Networks (DDQN), Conservative Q-Learning (CQL), and a behavior cloning (BC) model of clinician practice. CQL consistently exhibits stable learning dynamics and favorable Fitted Q Evaluation estimates, while DDQN is prone to overestimation and instability, particularly at finer temporal resolutions and with dense rewards. CQL policies achieve high action-level agreement with historical clinician decisions for both fluids and vasopressors and reproduce clinically plausible escalation patterns across sepsis severity strata, whereas DDQN policies diverge more frequently toward implausible dosing. Temporal aggregation emerges as a key regularizer: moving from 4-hour to 8-hour bins shortens horizons, smooths reward noise, and improves stability without erasing clinically meaningful dynamics, with 8-hour binning providing the best trade-off between policy performance and granularity. Our findings highlight time-step size as a core design choice in offline RL for healthcare and provide empirical evidence that alternatives beyond the conventional 4-hour setup can enhance stability and safety while preserving clinical interpretability.

Author summary

We studied how artificial intelligence might support treatment decisions for children with sepsis in intensive care. Sepsis is a serious condition that can worsen quickly, and clinicians often need to decide how much fluid or blood pressure support to give over time. Although artificial intelligence has been studied for adult sepsis, much less is known about how these methods perform in children, whose illness patterns and treatment needs can differ in important ways. Using records from 2,229 pediatric intensive care admissions, we tested whether a learning system could identify treatment strategies that were both stable and consistent with real clinical practice. We found that some modeling choices had a major effect on the safety and credibility of the recommendations. In particular, grouping data into 8-hour intervals produced more reliable results than the more commonly used 4-hour approach, while still preserving meaningful changes in illness severity. These results suggest that safe and useful decision-support tools for pediatric sepsis depend not only on the choice of the right algorithm, but also on careful design choices. We hope this work helps guide future research toward more trustworthy clinical artificial intelligence.

Article activity feed