Large Language Models Generate Stigmatizing Language During Reasoning Over Real-World Clinical Data
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Stigmatizing language in clinical documentation, which conveys negative stereotypes, attitudes, or judgments toward patients, is a recognized source of documentation bias and is associated with poorer care and adverse health outcomes. Although prior stigma-related research has explored on clinician-written EHR notes, the increasing use of large language model (LLM)-generated documentation in clinical workflows raises new concerns about its potential to produce or amplify bias and affect patient safety. In this study, we conducted a large-scale assessment of stigmatizing language in LLM-generated reasoning text on 35 real-world clinical tasks across 107 LLMs. We applied a psychiatrist-validated, natural language processing (NLP) system to detect stigma terms in LLM reasoning text and quantified stigma rates of LLM-generated reasoning texts across 3,745 model-task pairs. Results showed that stigma rates ranged from 0% to 33.33%, with 84.06% of pairs containing stigma terms. Reasoning models showed higher stigma rates than non-reasoning models (2.35% vs. 1.70%; p < 0.0001), whereas medical models did not show significantly lower stigma rates than general-purpose models (1.80% vs. 2.00%; p = 0.26). Stigma rates of LLM outputs correlated negatively with task accuracy (r = -0.283; p < 0.001) and positively with input clinical-text stigma (r = 0.569; p < 0.001), with 19.76% of LLM-task pairs amplifying stigma in the original input notes. The effectiveness of prompt engineering as a destigmatizing approach varied across models, with stigma rates reduced by up to 91.91% without compromising model performance. This study shows that stigmatizing language generation is common but modifiable in LLM-generated reasoning traces, highlighting the need for direct stigma-related safety evaluation and robust destigmatization strategies before clinical deployment.