Ensemble SHAP Aggregation and Attribution Variability in Clinical Machine Learning: A COVID-19 Mortality Study

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Objective

To combine performance-weighted ensemble SHAP aggregation with resampling-based assessment of feature-importance variability and patient-level attribution alignment for interpreting COVID-19 mortality predictions.

Methods

We analyzed a prospective cohort of 1,857 patients hospitalized with COVID-19 at two hospitals in Peru. Ten predictive algorithms were evaluated using five-fold cross-validation, and their mean AUROC values determined their attribution-aggregation weights. Within each model and training trial, signed SHAP values were normalized by the mean total absolute attribution across evaluation patients before performance-weighted aggregation. Feature importance was summarized across 30 resampled training trials using means and normal-approximation 95% confidence intervals. Patient-level attribution alignment was assessed using cosine similarity between signed feature-attribution vectors for corresponding patients, with whole-trial patient-correspondence randomization and Benjamini–Hochberg correction. Subgroup observations were reweighted toward the evaluated population’s feature distributions for secondary comparisons of existing absolute SHAP values.

Results

Among the 1,857 included patients, 982 (52.9%) died during hospitalization. Random forest achieved the highest mean AUROC (0.925 ± 0.010), followed by AdaBoost (0.919 ± 0.017) and logistic regression (0.918 ± 0.007), with variability reported as the SEM. Attribution profiles were more similar across repeated training trials of the same algorithm (mean Pearson r = 0.783; SEM, 0.004) than across algorithms within the same trial (mean Pearson r = 0.369; SEM, 0.006). The largest normalized attribution magnitudes were observed for dexamethasone use at home without oxygen support (6.186%; 95% CI, 5.795–6.577), PaO 2 /FiO 2 ratio (3.014%; 95% CI, 2.787–3.242), shortness of breath (2.774%; 95% CI, 2.592–2.956), and FiO 2 (2.560%; 95% CI, 2.284–2.836). Features differed in patient-level attribution alignment, indicating that importance magnitude and robustness provided complementary information. Among 1,763 eligible subgroup–feature comparisons, 64 had nominal one-sided p < 0.05, although none remained below 0.05 after Benjamini–Hochberg adjustment.

Conclusions

Algorithms with similar predictive performance produced substantially different feature-attribution profiles. Normalized performance-weighted SHAP aggregation provided a representative relative-importance summary across algorithms, while resampling-based alignment assessment qualified the consistency of individual feature contributions. The resulting attribution patterns describe fitted-model behavior and should not be interpreted as causal clinical effects.

Article activity feed