ICU Hour-24 Landmark Prediction of In-Hospital Mortality in Critically Ill Patients With Coronary Artery Disease: Development in MIMIC-IV and External Validation in eICU

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Retrospective intensive care unit prediction models are vulnerable to temporal leakage when predictors include information recorded after the intended prediction time. We developed and externally validated an ICU hour-24 landmark prediction pipeline for in-hospital mortality among critically ill patients with coronary artery disease using timestamp-restricted features from MIMIC-IV and eICU.

Methods

Adults with coronary artery disease or coronary heart disease who were alive and remained under ICU observation at hour 24 were included. MIMIC-IV was used for model development, and eICU was reserved for external validation. Dynamic events were restricted to ICU admission through hour 24 in MIMIC-IV and offsets of 0–1440 minutes in eICU before aggregation. We evaluated an XGBoost model using baseline and respiratory-support predictors and a 102-predictor random forest using baseline, respiratory-support, and treatment predictors. Robustness was assessed across 30 repeated patient-grouped MIMIC-IV validation splits, and external uncertainty was estimated using 1,000 subject-clustered bootstrap resamples.

Results

The MIMIC-IV cohort included 4,341 ICU stays with 993 deaths, and the eICU cohort included 19,464 stays with 2,237 deaths. In eICU, XGBoost achieved a ROC-AUC of 0.7973 (95% CI, 0.7878–0.8068), a PR-AUC (calculated as average precision) of 0.3627 (95% CI, 0.3423–0.3843), and a Brier score of 0.1189 (95% CI, 0.1164–0.1213). The random forest achieved a ROC-AUC of 0.8060 (95% CI, 0.7960–0.8154), a PR-AUC of 0.3687 (95% CI, 0.3471–0.3915), and a Brier score of 0.1032 (95% CI, 0.1012–0.1054). The random forest had a modestly higher ROC-AUC and lower Brier score than XGBoost. Mean validation ROC-AUCs across repeated MIMIC-IV splits were 0.7970 and 0.7954, respectively. Exploratory analyses suggested that narrower, consistently harmonized feature sets transported more reliably than broader expansions.

Conclusions

Timestamp-restricted first-day models achieved external ROC-AUCs of approximately 0.80. However, external calibration remained imperfect, and local recalibration and prospective evaluation would be required before clinical use.

Article activity feed