Offline Reinforcement Learning for Out-of-Distribution ICU Sepsis Decision Support

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Offline reinforcement learning (RL) provides a promising framework for learning and evaluating treatment policies from logged clinical data, particularly in sequential decision-making settings where prospective exploration would be unsafe. In ICU sepsis management, however, it remains unclear whether offline RL policies retain stable behavior under increasingly severe out-of-distribution (OOD) patient cohorts. In this paper, we evaluate standard offline RL methods on three severity-enriched OOD test mixtures from the MIMIC-III benchmark dataset to determine whether offline policies retain a stable, action-sensitive decision-support signal. Under the shared learned-dynamics off-policy evaluation (OPE) protocol, as the severe-OOD ratio increases from 25% to 75%, observed clinical survival declines from 67% to 49%, while the best offline method in each mixture receives model-predicted terminal survival values of 87%, 86%, and 85%, respectively. Because observed clinical survival and model-predicted terminal survival are different quantities, this contrast suggests a stable model-based decision-support signal under severity shift. We further present a secondary physiological stabilization analysis using an episode-level physiological stabilization score (EPSS), a heuristic summary of whether selected physiological variables move in favorable directions during follow-up. In this analysis, model-generated rollouts under offline policies receive higher EPSS values than matched logged clinical trajectories for several physiological components. Together, these results support learned-dynamics OPE as a useful severity-OOD stress test for offline RL policies in ICU sepsis, while leaving prospective and causal validation as necessary next steps.

Article activity feed