Are automated documentation-error judges fit to measure ambient AI scribes? A pre-registered, blinded human-validation study

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Objectives

Safety claims for ambient artificial intelligence (AI) scribes rest on automated judges that detect documentation errors and grade clinical risk. Expert reviewers are under-sensitive and disagree with one another, so no gold standard exists and validation cannot mean accuracy. We tested whether such judges are a defensible instrument: reproducible, within the envelope of expert disagreement, and non-differential across arms.

Methods

Pre-registered, blinded validation study nested in a multi-country simulation of ambient AI documentation (English setting), reported per GRRAS. Ten external clinicians independently adjudicated a stratified sample of 434 pipeline flags, retained and screen-discarded, blinded to note authorship, identification source, the pipeline’s verdict and severity tier. Agreement used Gwet’s AC1; proportions carry Wilson intervals. Three propositions were pre-specified: envelope parity, non-differential behaviour across arms, and concordance on consensus cases.

Results

All ten reviewers completed: 565 adjudications across 434 items, 131 of them double-rated. Inter-clinician agreement on genuineness was fair (raw 59%, 95% CI 50 to 67; AC1 0.24), leaving no human consensus to serve as truth. Judge-clinician agreement was 64% (95% CI 60 to 68), overlapping that interval. Behaviour was near-symmetric on contrast-critical metrics: kept-precision 74% for AI against 81% for clinician notes, and severity signed gap +0.06 against -0.09 tiers. One sub-metric was asymmetric: removed-confirmed 56% against 42%, so the screen over-removes more on clinician notes, a direction conservative to the parent contrast. On 77 consensus items the pipeline concurred on 70% (95% CI 59 to 79). Latent-class triangulation placed the genuine-error rate among flagged candidates at 68% (94% credible interval 48 to 83).

Conclusions

The judges behave as a consistent, near-non-differential, clinician-equivalent instrument. This licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions. It is not a claim of accuracy, which moderate consensus concordance and fair reliability preclude, and the genuine-error rate is best reported as an interval.

What is already known on this topic

  • Ambient AI scribes are entering routine practice, and their evaluation depends on detecting documentation errors at a scale that unaided human review cannot achieve

  • Large language models have been validated as judges of clinical text quality, with strong agreement against human raters on ordinal quality instruments

  • Human record review is known to be under-sensitive, and expert reviewers agree only moderately even on explicit quality criteria, so no gold standard for documentation error exists

What this study adds

  • Clinicians agreed only fairly on whether a flagged item was a genuine documentation error (AC1 0.24), which is substantially weaker than published agreement on ordinal note-quality scoring and indicates that error detection is the harder judgement.

  • An automated judge agreed with clinicians about as well as clinicians agreed with one another, and behaved near-symmetrically across AI-authored and clinician-authored notes on the metrics that determine a comparative contrast.

  • The one asymmetry detected runs against the sponsor’s interest, and the genuine-error rate among flagged candidates is identifiable only as an interval.

How this study might affect research, practice or policy

  • Where no gold standard exists, validation of automated documentation-error judges should be framed as defensibility, meaning reproducibility, parity with the expert envelope and non-differential behaviour, in place of accuracy.

  • Comparative studies of AI and clinician documentation should report the direction in which any measurement asymmetry biases their own contrast.

  • Absolute documentation-error rates derived from automated judges should be reported as bounded intervals.

Article activity feed