Accuracy and error patterns of ChatGPT-4o for real-time English–Nepali voice translation: A cross-sectional field evaluation in rural Nepal
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems (AI) can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English–Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal.
In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali–English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations.
Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 ± 0.53 vs 2.23 ± 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker.
ChatGPT-4o demonstrated potential for real-time English–Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.
Author Summary
Being able to understand and be understood is essential in health care and community research. Yet trained interpreters are not always available, particularly in rural or resource-limited settings. Voice-enabled artificial intelligence could help fill some communication gaps, but a translation can sound convincing without accurately representing what a person said.
We examined how ChatGPT-4o performed during real English–Nepali conversations in a community setting in Nepal. Rather than testing carefully prepared written sentences, we studied spoken exchanges that included natural pauses, varied responses, background conditions, and culturally specific ways of speaking. We found that the tool often communicated the basic message, but it also changed meanings, omitted details, added information, and sometimes used language that sounded unnatural to Nepali speakers. Importantly, some inaccurate translations remained fluent enough that an English-speaking listener might not recognize the error.
Our study shows why natural-sounding artificial intelligence should not automatically be considered reliable. These tools may assist with selected low-risk conversations when other language support is unavailable, but they should not replace trained interpreters or bilingual review when misunderstandings could affect people’s health or representation in research.