Auditing Large Language Model–Generated Digital Standardized Patients for Demographic Bias: A Simulation Study with HIV Pre-Exposure Prophylaxis Screening as a Tracer Condition
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Large language models (LLMs) are entering clinical training as digital standardized patients (DSPs), simulated patient encounters the model scripts and portrays. Demographic associations learned from corpus co-occurrences can enter at two points: the written case scripts and the live portrayals improvised in role-plays. We audited both pathways using an HIV pre-exposure prophylaxis (PrEP) screening scenario across six demographic factors (age, gender, marital status, sexual orientation, education, and race or ethnicity) in a 216-cell factorial design. With a generated arm and a template-substituted control arm, we simulated 4,320 conversations with fixed audit questions, measuring three channels: the case scripts, the composite role-play trainees receive, and role-play under identical control cases. Differences concentrated where corpus associations were strongest and reinforced predictable stereotypes. Anal sex appeared in 100% of gay men’s cases, 83% of bisexual men’s and 3% of heterosexual men’s; the six factors explained 32% of the variance in composite sexual risk; and education predicted assigned socioeconomic status ( R̄ 2 = 0.59). Demographics predicted 5 of the 9 improvised probe responses in composite role-play, and 2 under control cases. These two pathways sometimes had opposite signed effects: the model wrote women with higher alcohol use, but role-played them as drinking less, and the two canceled in composite role-play. An audit of either pathway alone would have obscured both effects. These issues are fixable as they are predictable and consistent. LLM-generated DSPs must have their cases and live portrayals audited as separate objects before deployment.
Author summary
Medical schools are beginning to use artificial intelligence chatbots as practice patients: the software writes a patient’s story and then plays that patient in a conversation with a trainee. We asked whether these simulated patients portray people differently depending on their demographic profile. We built 216 patient profiles that differed only in age, gender, marital status, sexual orientation, education, and race, asked a widely used language model to write a case for each one in a scenario about HIV prevention, and then ran thousands of practice conversations with the model playing the patient. The written cases assigned riskier behavior to some groups: nearly every case describing a gay man mentioned anal sex, and a profile’s education level predicted the social class the model invented for it. The live conversations showed a second, separate pattern, sometimes in the opposite direction, so the two distortions could hide each other when combined. A program that checks only the written case, or only the conversation, would miss half the picture. We recommend that training programs audit both before students meet these simulated patients.