Consumer chatbots gave similarly empathic answers whether safe or unsafe: a physician-rated evaluation in six languages

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Patients cannot check the clinical content of chatbot health advice, so they judge it by what they can perceive. We examined whether physician-rated empathy tracked clinical quality, and whether that relationship held when the question was asked in another language.

Methods

Four consumer chatbots answered forum-derived, clinician-adapted patient scenarios in six languages (English, Hebrew, French, Russian, Arabic, Thai), yielding 504 responses. Two language-matched physicians per language, blinded to chatbot identity, rated accuracy, safety, referral, cultural appropriateness, and empathy, and completed an item-level checklist, giving 1,008 ratings across 21 scenarios. Associations were estimated within physicians. The same two physicians rated English and Hebrew, so that contrast was also within-physician.

Results

Empathy did not separate safe from unsafe responses (AUC 0.49, 95% CI 0.39 to 0.62), and within-physician slopes on safety and substance were near zero (−0.006 and −0.004). Each dimension correlated with its own checklist (r = 0.81 and 0.70) and not with the other (0.01 and −0.01). The substance-minus-empathy gap narrowed from 0.92 to 0.44 and from 1.09 to 0.52 in the two English– Hebrew physicians, driven by lower substance. Unsafe ratings concentrated on the same three scenarios across products (p<0.001), and ten responses were accurate yet unsafe.

Conclusions

Empathy, the one cue a patient can judge, carried no information about whether the advice was safe, and clinical substance fell in Hebrew within both physicians who rated it. Evaluation should score clinical content independently of empathy, in each deployment language, and anchor on high-risk scenarios rather than any single product.

Article activity feed