Using explainable machine learning to characterise data drift and detect emergent health risks for emergency department admissions during COVID-19

Christopher Duckworth
Francis P. Chmiel
Dan K. Burns
Zlatko D. Zlatev
Neil M. White
Thomas W. V. Daniels
Michael Kiuber
Michael J. Boniface

This article has been Reviewed by the following groups

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

Evaluated articles (ScreenIT)

Abstract

A key task of emergency departments is to promptly identify patients who require hospital admission. Early identification ensures patient safety and aids organisational planning. Supervised machine learning algorithms can use data describing historical episodes to make ahead-of-time predictions of clinical outcomes. Despite this, clinical settings are dynamic environments and the underlying data distributions characterising episodes can change with time (data drift), and so can the relationship between episode characteristics and associated clinical outcomes (concept drift). Practically this means deployed algorithms must be monitored to ensure their safety. We demonstrate how explainable machine learning can be used to monitor data drift, using the COVID-19 pandemic as a severe example. We present a machine learning classifier trained using (pre-COVID-19) data, to identify patients at high risk of admission during an emergency department attendance. We then evaluate our model’s performance on attendances occurring pre-pandemic (AUROC of 0.856 with 95%CI [0.852, 0.859]) and during the COVID-19 pandemic (AUROC of 0.826 with 95%CI [0.814, 0.837]). We demonstrate two benefits of explainable machine learning (SHAP) for models deployed in healthcare settings: (1) By tracking the variation in a feature’s SHAP value relative to its global importance, a complimentary measure of data drift is found which highlights the need to retrain a predictive model. (2) By observing the relative changes in feature importance emergent health risks can be identified.

Version published to 10.1038/s41598-021-02481-y
Nov 26, 2021
ScreenIT
Jun 3, 2021
SciScore for 10.1101/2021.05.27.21257713: (What is this?)
Please note, not all rigor criteria are appropriate for all manuscripts.
Table 1: Rigor
NIH rigor criteria are not applicable to paper type.
Table 2: Resources
No key resources detected.
Results from OddPub: We did not detect open data. We also did not detect open code. Researchers are encouraged to share open data when possible (see Nature blog).
Results from LimitationRecognizer: We detected the following sentences addressing limitations in the study:
One limitation of this study is that we only possess prior patient information from previous discharge summaries at UHS, and, hence only have patient histories for 55% of our recorded attendances. Including patient histories from their non-emergency and routine treatment would enable even more predictive early warning modelling.
Results …
SciScore for 10.1101/2021.05.27.21257713: (What is this?)
Please note, not all rigor criteria are appropriate for all manuscripts.
Table 1: Rigor
NIH rigor criteria are not applicable to paper type.
Table 2: Resources
No key resources detected.
Results from OddPub: We did not detect open data. We also did not detect open code. Researchers are encouraged to share open data when possible (see Nature blog).
Results from LimitationRecognizer: We detected the following sentences addressing limitations in the study:
One limitation of this study is that we only possess prior patient information from previous discharge summaries at UHS, and, hence only have patient histories for 55% of our recorded attendances. Including patient histories from their non-emergency and routine treatment would enable even more predictive early warning modelling.
Results from TrialIdentifier: No clinical trial numbers were referenced.
Results from Barzooka: We did not find any issues relating to the usage of bar graphs.
Results from JetFighter: We did not find any issues relating to colormaps.
Results from rtransparent:
Thank you for including a conflict of interest statement. Authors are encouraged to include this statement when submitting to a journal.
Thank you for including a funding statement. Authors are encouraged to include this statement when submitting to a journal.
No protocol registration statement was detected.
Results from scite Reference Check: We found no unreliable references.
About SciScore
SciScore is an automated tool that is designed to assist expert reviewers by finding and presenting formulaic information scattered throughout a paper in a standard, easy to digest format. SciScore checks for the presence and correctness of RRIDs (research resource identifiers), and for rigor criteria such as sex and investigator blinding. For details on the theoretical underpinning of rigor criteria and the tools shown here, including references cited, please follow this link.
Read the original source
Version published to 10.1101/2021.05.27.21257713 on medRxiv
May 29, 2021

Detection and Early Warning for Patient's Critical Condition Using Bayes’ Classifier

This article has 5 authors:
1. Shuo-Tsung Chen
2. Shan-Ju Lin
3. Chur-Jen Chen
4. Shu-Yi Tu
5. Hao-Chun Lu
This article has no evaluationsLatest version Dec 1, 2025
A Prospective Real-time Early Warning System to Anticipate Onsets and Peaks of Respiratory Diseases Outbreaks at the State Level in the U.S. A Transfer Learning Approach Leveraging Digital Traces

This article has 6 authors:
1. Raul Garrido Garcia
2. Leonardo Clemente
3. Austin Meyer
4. George Dewey
5. Shihao Yang
6. Mauricio Santillana
This article has no evaluationsLatest version Oct 13, 2025
Building a Machine Learning Model to Predict the Early Mortality Risk in Pediatric ICU Sepsis Patients

This article has 6 authors:
1. Lin Yang
2. Na Zang
3. Ying Yang
4. KaiBing Pu
5. Cong Liu
6. LiPing Tan
This article has no evaluationsLatest version Nov 3, 2025

This article has been Reviewed by the following groups

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Detection and Early Warning for Patient's Critical Condition Using Bayes’ Classifier

A Prospective Real-time Early Warning System to Anticipate Onsets and Peaks of Respiratory Diseases Outbreaks at the State Level in the U.S. A Transfer Learning Approach Leveraging Digital Traces

Building a Machine Learning Model to Predict the Early Mortality Risk in Pediatric ICU Sepsis Patients