“Mortality Prediction in Heart Failure using Explainable AI: Development and Validation of the EXACT-HF Model in a Nationwide Multi-Ethnic Asian Cohort”

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Established heart-failure (HF) risk scores were developed largely in Western cohorts, assume linear predictor effects, and stop at prognosis without indicating action. We developed and externally validated machine-learning (ML) models for 30-day and 1-year mortality in a multi-ethnic Asian HF population, benchmarked against three established scores, and tested whether they localised correctable guideline-directed medical therapy (GDMT) gaps.

Methods

We analysed 9,050 adults hospitalised with HF across eight Singapore hospitals (November 2016–March 2023). Gradient-boosted trees (XGBoost, primary) and elastic-net logistic regression were developed on three hospitals (n=4,034) and validated on five independent hospitals (n=5,016), without patient or site overlap. Comparators (MAGGIC, Singapore HF Score, OPTIMIZE-HF-Asia) were logistically recalibrated on development data.

Results

On external validation XGBoost achieved AUROCs of 0.862 (95% CI 0.833–0.890) for 30-day and 0.749 (0.734–0.764) for 1-year mortality, versus 0.711–0.754 and 0.682–0.699 for the three recalibrated scores; logistic regression performed near-identically (0.864, 0.757) and was better calibrated at 1 year. Superiority held under temporal validation, multiple imputation and optimism correction, although the 30-day advantage was no longer distinguishable from MAGGIC once medications were removed. Among predicted-high-risk patients with reduced ejection fraction, mean GDMT exposure was 0.92 of four pillars versus 1.87 across all HFrEF patients, and 36.3% received none of the four versus 2.0% of low-risk patients. Ten variables reproduced full-model 30-day discrimination (AUROC 0.865).

Conclusions

In external validation, interpretable models substantially outperformed established HF risk scores for short- and long-term mortality and simultaneously localised correctable therapeutic gaps, supporting a prediction-to-action tool deployable with ten bedside variables.

Clinical Perspective

What Is New?

  • In 9,050 patients hospitalised for heart failure across eight Singapore hospitals, prediction models estimated 30-day and 1-year all-cause mortality substantially better than three established risk scores (30-day AUROC 0.86 versus 0.71–0.75), validated in five hospitals independent of model development.

  • Ten routinely recorded bedside variables reproduced the full 85-variable model’s 30-day discrimination. This enables bedside risk estimation without electronic health record integration.

What Are the Clinical Implications?

  • The same model identifies both who is at risk and which guideline-directed therapies they lack: flagged patients were receiving a mean of 0.9 of the four pillars, against 1.9 across all patients with reduced ejection fraction, although most non-prescription had a documented renal or haemodynamic contraindication, so this correctable fraction is an upper bound requiring prospective validation.

Article activity feed