“Mortality Prediction in Heart Failure using Explainable AI: Development and Validation of the EXACT-HF Model in a Nationwide Multi-Ethnic Asian Cohort”
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
Established heart-failure (HF) risk scores were developed largely in Western cohorts, assume linear predictor effects, and stop at prognosis without indicating action. We developed and externally validated machine-learning (ML) models for 30-day and 1-year mortality in a multi-ethnic Asian HF population, benchmarked against three established scores, and tested whether they localised correctable guideline-directed medical therapy (GDMT) gaps.
Methods
We analysed 9,050 adults hospitalised with HF across eight Singapore hospitals (November 2016–March 2023). Gradient-boosted trees (XGBoost, primary) and elastic-net logistic regression were developed on three hospitals (n=4,034) and validated on five independent hospitals (n=5,016), without patient or site overlap. Comparators (MAGGIC, Singapore HF Score, OPTIMIZE-HF-Asia) were logistically recalibrated on development data.
Results
On external validation XGBoost achieved AUROCs of 0.862 (95% CI 0.833–0.890) for 30-day and 0.749 (0.734–0.764) for 1-year mortality, versus 0.711–0.754 and 0.682–0.699 for the three recalibrated scores; logistic regression performed near-identically (0.864, 0.757) and was better calibrated at 1 year. Superiority held under temporal validation, multiple imputation and optimism correction, although the 30-day advantage was no longer distinguishable from MAGGIC once medications were removed. Among predicted-high-risk patients with reduced ejection fraction, mean GDMT exposure was 0.92 of four pillars versus 1.87 across all HFrEF patients, and 36.3% received none of the four versus 2.0% of low-risk patients. Ten variables reproduced full-model 30-day discrimination (AUROC 0.865).
Conclusions
In external validation, interpretable models substantially outperformed established HF risk scores for short- and long-term mortality and simultaneously localised correctable therapeutic gaps, supporting a prediction-to-action tool deployable with ten bedside variables.
Clinical Perspective
What Is New?
-
In 9,050 patients hospitalised for heart failure across eight Singapore hospitals, prediction models estimated 30-day and 1-year all-cause mortality substantially better than three established risk scores (30-day AUROC 0.86 versus 0.71–0.75), validated in five hospitals independent of model development.
-
Ten routinely recorded bedside variables reproduced the full 85-variable model’s 30-day discrimination. This enables bedside risk estimation without electronic health record integration.
What Are the Clinical Implications?
-
The same model identifies both who is at risk and which guideline-directed therapies they lack: flagged patients were receiving a mean of 0.9 of the four pillars, against 1.9 across all patients with reduced ejection fraction, although most non-prescription had a documented renal or haemodynamic contraindication, so this correctable fraction is an upper bound requiring prospective validation.