A 3-Minute Education on the False Positive Paradox Improves Trust Calibration in AI-Assisted Intracranial Aneurysm Detection: A Multinational Randomized Controlled Reader Study

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Even a highly accurate diagnostic test can yield more false-positive than true-positive findings in low-prevalence settings, which is known as the false positive paradox. Radiologists’ unawareness of this paradox may foster automation bias, the tendency to excessively rely on AI outputs.

Methods

In this prospective, multinational, randomized controlled reader study (DRKS00038740), 34 readers from 10 countries (16 residents, 8 general radiologists or fellows, and 10 neuroradiologists) were randomly assigned to a control group (n = 17) or intervention group (n = 17), stratified by experience level. The intervention group reviewed a short, 3-minute educational video explaining the false positive paradox prior to the reading session. Both groups evaluated 20 TOF-MRA studies with AI-flagged findings (10% true-positive, 90% false-positive). Primary outcomes were acceptance rate of false-positive AI findings and follow-up intensity. These were evaluated using mixed models with crossed random effects for reader and case.

Results

At baseline, readers vastly overestimated the positive predictive value of AI tools for intracranial aneurysm detection (mean estimate, 62.9%; simulation-based estimate, 15.4% [95% interval, 8.1-28.0%]). The intervention reduced the odds of accepting AI false positives (OR 0.50 [upper 95% confidence bound, 0.95], one-sided p = 0.017), with acceptance probabilities of 12.7% (95% CI, 6.0-25.0%) in the intervention group compared to 22.5% (95% CI, 11.6-39.2%) in the control group. The intervention group exhibited a downward shift in follow-up intensity for false positives (OR 0.47 [upper 95% confidence bound, 0.81]; one-sided p = 0.014), recommending follow-up in 39.2% (120/306) of cases, compared to 54.9% (168/306) in the control group.

Conclusion

A brief education on the false positive paradox improved trust calibration in AI-assisted intracranial aneurysm detection. Our findings highlight the potential of reader-side cognitive debiasing strategies to improve trust calibration and support safer use of AI in radiology.

Summary

A brief education on realistic positive predictive value ranges reduced radiologists’ uncritical acceptance of false-positive AI flags for intracranial aneurysm and the intensity of recommended follow-up.

Key Results

  • -

    In a multinational, randomized controlled reader study with 34 radiologists from 10 countries, readers initially vastly overestimated the positive predictive value of AI for intracranial aneurysm detection (mean estimate, 62.9%; simulation-based estimate, 15.4% [95% interval, 8.1-28.0%]), demonstrating base-rate neglect.

  • -

    Exposure to a 3-minute educational video explaining the false positive paradox reduced the odds of accepting false-positive AI findings (OR 0.50 [upper 95% confidence bound 0.95], one-sided p = 0.017), with an acceptance probability of 12.7% (95% CI, 6.0-25.0%) in the intervention group compared to 22.5% (95% CI, 11.6-39.2%) in the control group.

  • -

    The intervention group had lower odds of recommending a more intensive follow-up strategy (OR 0.47 [upper 95% confidence bound, 0.81]; one-sided p = 0.014).

  • Article activity feed