Reconsidering the case against risk prediction in self-harm: routinely collected health data distinguishes groups at higher and lower risk of adverse outcomes following paracetamol overdose

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

UK clinical guidance recommends that structured risk prediction tools and risk stratification should not be used in self-harm, to predict suicide or determine who is offered treatment. Underpinning this position is the premise that routinely collected health data contain no useful predictive signal, which has received little direct scrutiny.

Objective

To test whether routinely collected electronic health record data can distinguish groups at higher and lower risk of severe outcomes following paracetamol overdose.

Methods

We analysed 4,095 adults presenting to NHS Lothian emergency departments with paracetamol overdose (2017-2023). Elastic-net logistic regression was fitted to 37 routinely collected electronic health record features to predict a composite of death or mental health inpatient admission at 0-7, 8-30 and 31-365 days following attendance, evaluated on a held-out 20% test set with bootstrapping.

Findings

Events occurred in 5.5% of patients at 0-7 days, 2.0% at 8-30 days and 7.9% at 31-365 days, dominated by mental health admission. Bootstrap AUROC 95% confidence intervals lay above 0.5 in every window (0.65-0.82, 0.63-0.90, 0.71-0.85): models ranked patients better than chance. Calibration slopes (1.04, 1.14, 1.07) were close to one. Ranking drew primarily on mental health-related features.

Conclusions

Routinely collected health data carried predictive signal for severe outcomes after paracetamol overdose, although discrimination fell short of what is needed for individual-level clinical use.

Clinical implications

These models are not proposed for clinical deployment; however, treating risk prediction as a settled question will redirect research efforts, potentially excluding this patient population from machine learning advances driving improvements in care in other medical specialties.

Summary box

What is already known on this topic

NICE NG225 (2022), NHS England’s Staying Safe from Suicide framework (2025) and NCISH guidance (2024) recommend against structured risk prediction in self-harm, both for predicting suicide or repetition and for allocating treatment. An increasingly common reading of the underpinning literature is that routinely collected health data contain no useful predictive signal in this population. Evaluations of this claim in large-scale UK linked healthcare datasets remain scarce.

What this study adds

In a whole-population cohort of adults attending emergency departments with paracetamol overdose, models using 37 structured electronic health record features separated groups at higher and lower risk of all-cause mortality or mental health admission across all three time horizons, with calibration slopes close to one. Discriminative signal was carried primarily by mental health-related features rather than demographics. Useful predictive signal is therefore recoverable from routinely collected UK healthcare data in this population.

How this study might affect research, practice or policy

These findings are not evidence that suicide can be predicted at the individual level, and the models are not proposed for clinical use. They do, however, indicate that the empirical question underlying current guidance remains open, and should stay subject to scientific enquiry and to revision if sufficiently robust models emerge. Future research should expand the feature space beyond structured records, particularly into clinical free text, where the information on which psychosocial assessment relies already resides.

Article activity feed