Automating clinical trial outcome identification and misreporting detection using RegCheck
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Objective
To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project.
Design
Validation study.
Setting
Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals.
Participants
Published clinical trial reports and their corresponding prespecified registrations and/or protocols.
Main outcome measures
Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow.
Results
Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck’s overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck’s judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper.
Conclusions
RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.
What is already known on this topic
Undeclared discrepancies between prespecified and reported trial outcomes remain common and can distort the medical evidence base.
Manual auditing projects such as COMPare Trials have shown that outcome switching is prevalent, but this work is labour intensive and difficult to sustain at scale.
Large language models may help automate registration-to-publication comparisons, but their performance against established manual benchmarks has not been well characterised.
What this study adds
RegCheck, an LLM-based workflow, achieved high performance for outcome extraction, generally high performance for outcome classification (with three protocols exhibiting exceptionally poor performance), and high performance for misreporting detection when benchmarked against COMPare Trials.
Manual adjudication of disagreements in misreporting judgements suggested that RegCheck’s judgements of misreporting were (at times) more accurate than those of the human reports.
The workflow operated at a low mean cost per paper, suggesting that automated trial checking could be deployed at scale as a screening and decision-support tool.