Do established comorbidity scores predict in-hospital mortality equally well in women and men? A nationwide validation in 166.7 million German inpatient cases, 2010–2024

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Objectives

To determine whether three established comorbidity scores (Charlson Comorbidity Index, Elixhauser Comorbidity Sum, van Walraven Score), applied with a single formula to both sexes, are equally well calibrated and equally discriminating in women and men, and whether any difference reaches a pre-specified threshold of clinical relevance.

Design

Population-based retrospective cohort study; external validation of three pre-existing prediction models (TRIPOD Type 4) with sex-stratified evaluation. The analysis plan was fixed in the time-stamped Research Data Centre submission syntax before any aggregated data existed.

Setting

Germany; complete national enumeration of acute somatic Diagnosis-Related Group (DRG)-billed adult inpatient cases, reporting years 2010 to 2024.

Participants

166 693 491 adult inpatient cases meeting a pre-specified nine-step eligibility filter.

Main outcome measures

In-hospital mortality. Primary axis: calibration (calibration-in-the-large, CITL; calibration slope; observed mortality per score quintile). Secondary axis: discrimination (area under the receiver operating characteristic curve, AUC). Pre-specified thresholds: between-sex slope difference of at least 0.05, absolute AUC difference of at least 0.01.

Results

Calibration-in-the-large was effectively perfect in both sexes (|CITL| ≤ 0.0004). The scores diverged in how prediction kept pace with rising risk, and in opposite directions for women and men: for the Charlson index the calibration slope was 0.968 in men and 1.034 in women (difference 0.066), with smaller differences for the Elixhauser-Sum (0.039) and the van Walraven Score (0.029). At the same score quintile, observed mortality was higher in men at most quintiles, with a maximum risk ratio of 1.21 (95 % CI 1.20 to 1.22); two van Walraven quintiles ran the other way (0.87 and 0.97). Differences in discrimination were small (random-effects AUC difference: Charlson −0.0076, 95 % CI −0.0095 to −0.0057; Elixhauser-Sum +0.0074; van Walraven −0.0026).

Conclusions

Applied with a single formula to both sexes, three established comorbidity scores are calibrated differently in women and men, whereas their discrimination differs only marginally. Because these scores are reused at scale in research and in risk-adjusted mortality comparisons between hospitals, a systematic difference in calibration does not average out. This supports sex-specific recalibration. The downstream effect on hospital benchmarking is not yet quantified and should not be presumed negligible.

Study registration

OSF DOI 10.17605/OSF.IO/P3QAW.

Plain Language Summary

Comorbidity scores allow for cross-hospital mortality comparisons for how ill patients already were on admission, by combining their pre-existing conditions into a single risk number. Worldwide, three scores are predominantly used: The Charlson Comorbidity Index, the Elixhauser Comorbidity Sum, and the van Walraven Score. All three are applied to men and women using the same formula. Only the van Walraven Score was originally built to predict death in hospital. The Charlson index was built for long-term survival, and the Elixhauser Sum is a simple unweighted count of conditions. None had been checked separately for men and women in a large population.

We applied the three scores to 166.7 million adult hospital stays in Germany between 2010 and 2024, selected from the national hospital record by pre-specified inclusion criteria. Separately for men and women, we analysed two things: Whether the predicted risk matched the mortality actually observed (calibration), and how well the score separated patients who died in hospital from those who survived (discrimination). Calibration differed systematically between the sexes. At the same score, men with the same predicted risk in fact died more often than women, by up to about a fifth across risk groups, and the scores tracked rising risk at slightly different rates in men and women, in opposite directions. Differences in how well the scores told apart those who died from those who survived were very small between the sexes, smaller than a thousandth on a scale that runs from 0.5 to 1.

For an individual patient these differences are too small to change a treatment decision. They matter because these scores are used everywhere, to compare the mortality of whole hospitals and in thousands of research studies, and a score that fits one sex better than the other bends those comparisons in a consistent direction. A small bias that always points the same way, repeated across millions of cases, does not cancel out. This is why the calibration differences, not the very small discrimination differences, are the result that counts, and why their full effect still needs to be measured. The pattern held across all 15 years and across age groups.

What this paper adds

  • Comorbidity scores derived from hospital discharge data are applied to men and women with a single single formula for both sexes, although sex-related differences in calibration of the Charlson index has been reported for long-term outcomes. No nationwide evaluation of all three of the most widely used scores by sex existed for in-hospital mortality.

  • In a sex-stratified validation on 166 693 491 German adult inpatient cases (2010– 2024), all three scores were systematically miscalibrated by sex when applied with a single formula to both sexes: At equal scores men had higher in-hospital mortality than women (risk ratios up to 1.21), while discrimination differed only marginally. The pattern held across all 15 years and across the age range, and reversed between emergency and elective admissions.

  • Because these scores are reused at scale in hospital benchmarking and observational research, a systematic sex bias in their calibration does not average out but is reproduced wherever they are applied with a single formula to both sexes. This justifies investigating sex-specific recalibration, and the downstream impact should not be presumed negligible.

Article activity feed