An interpretable Day-1 machine-learning model for COVID-19 mortality and severity prognosis: development in a Pakistani hospital cohort and external evaluation in Chinese and Italian populations
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Early risk stratification of COVID-19 patients at admission informs triage and resource allocation, particularly in low- and middle-income settings where intensive-care capacity is constrained. Most published models are not interpretable, rarely externally validated, and developed almost exclusively in high-income cohorts. We developed an interpretable random-forest classifier for three-class COVID-19 severity (Mild: not ventilated, survived; Severe: ventilated, survived; Fatal: died) using only data available within 24 hours of admission in 321 hospitalised PCR-confirmed COVID-19 patients in Pakistan. From 2,278 raw variables we retained 20 predictors by clinical screening and permutation-importance selection, with leakage-controlled multiple imputation by chained equations and SHapley Additive exPlanations (SHAP) for interpretation. A 15-feature variant was evaluated in two independent public cohorts: the Chinese iCTCF cohort (n = 894) and an Italian complete-blood-count cohort (n = 1,218). The development model achieved a macro F1 of 0.55, identifying the Fatal class most reliably (F1 0.67, recall 0.73) and Mild cases least well. SHAP identified a parsimonious set of admission-time predictors (oxygen saturation, lactate dehydrogenase, urea, C-reactive protein, age and radiographic pneumonia), including a raised urea-to-creatinine ratio in fatal cases most consistent with dehydration (median urea 80 vs 50 mg/dL, fatal vs mild; p < 0.001). Externally, the model retained useful discrimination in the Chinese cohort (mortality AUC 0.82 in 719 patients with recorded outcomes; severe-or-worse AUC 0.75; macro F1 0.50) but only modest discrimination in the Italian cohort (AUC 0.62, 95% CI 0.56-0.68), which we treat as a feature-availability stress test rather than a second validation. Cross-cohort SHAP rankings correlated strongly in China (Spearman ρ = 0.91) and moderately in Italy (ρ = 0.76), where most laboratory predictors were imputed and carried no information beyond the complete blood count; discrimination tracked how many predictors were actually measured. The admission-time signature converges with established markers of hypoperfusion and inflammation.