Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Objective

Ambient AI scribes evaluated using frequency-based metrics such as word error rate (WER), which do not represent clinical consequence. We tested whether variation in these metrics tracks consequential transcription errors.

Methods

We constructed a multilingual corpus from five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe. Six frequency metrics were computed. Three independent large language model raters from external providers assessed clinically meaningful error patterns in context using a Severity x Likelihood framework informed by UK digital clinical-safety-risk-management principles.

Results

Across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. None of six frequency metrics showed a statistically detectable association with serious clinical risk across languages; correlations were small (absolute Spearman ρ<0.16). A Severity × Likelihood sum remained strongly correlated with WER (ρ=0.80), showing that the aggregate remained dominated by benign errors. At complexity level 3, low-resource languages had worse WER than high-resource languages (β=+0.078, 95% CI +0.045 to +0.111; p<0.0001), without a detectable difference in CRITICAL/HIGH risk (OR 1.21, 95% CI 0.43 to 3.43; p=0.72).

Consultation complexity was the principal predictor of serious risk (OR 3.06 per level, p<0.0001).

Conclusion

Across this controlled multilingual corpus, aggregate transcription-frequency metrics did not reliably track the sparse severe tail of clinically consequential errors. WER remains appropriate for transcription quality, but these data do not support its use alone as a proxy for clinical safety.

Context-aware assessment of error consequence provides complementary information that frequency measures can dilute.

Highlights

  • WER did not track serious clinical-risks rate across 77 languages

  • Five related frequency metrics showed the same dissociation

  • Severity-weighting remained dominated by common low-risk errors

  • Quality tracked language resource; serious risk tracked complexity

  • We provide a reusable context-aware clinical-risk instrument

Article activity feed