Towards Interpretable Risk: Multidimensional Context for ICU Mortality Predictions

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

ICU mortality models can achieve strong discrimination, yet a risk score alone provides limited context for patient-level interpretation. We developed a multidimensional prediction-context framework that complements a calibrated mortality estimate with model behavior, data availability, recent physiology, and model attribution. Using 3,236 held-out ICU episodes from the MIMIC-III in-hospital mortality benchmark, we compared LSTM, GRU-D, XGBoost, and a weighted ensemble. The ensemble achieved an AUROC of 0.871 and AUPRC of 0.536; validation-based logistic recalibration improved the Brier score from 0.134 to 0.075. Incorrect predictions showed greater component-model disagreement and smaller decision-boundary margins, although model agreement and large margins did not guarantee correctness. Observation coverage and recent physiological trends also varied substantially across patients, highlighting differences in the information surrounding otherwise similar risk estimates. SHAP analysis attributed 89.9% of total absolute XGBoost attribution to physiological-value features and 10.1% to observation-process features. These dimensions were integrated into patient-level profiles to provide a more complete view of how predictions were formed and the clinical and data context surrounding them, extending interpretation beyond risk scores and feature rankings alone. Abbreviations: ICU (Intensive care unit)

Article activity feed