Do Large Language Models Use the Clinical Vignette? A Question-Ablation Study on the Orthopaedic In-Training Examination
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relative contributions of the clinical vignette, imaging, and answer options to model performance. We evaluated the contribution of these question components to LLM performance on the Orthopaedic In-Training Examination (OITE).
Methods
We evaluated three open-source Ministral-3 models (3B, 8B, and 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, and GPT-5.6 Terra) using 792 OITE questions from 2020 through 2024, comprising 434 questions containing clinical images and 358 without images. Question components were sequentially removed across five image-containing and three non–image-containing conditions. Accuracy was summarized using the median and IQR and compared using paired analyses with Holm adjustment. Semantic similarity between model-generated explanations and reference discussions was assessed using BioMedBERT.
Results
For image-containing questions, pooled accuracy was 59.13% (58.21–60.08) with complete information and 59.13% (58.18-60.11) without images. Accuracy decreased to 49.05% (48.07– 50.00) without the clinical vignette, 45.68% (44.70–46.69) without the vignette and images, and 36.49% (35.51–37.44) when the vignette, images, and question were removed. For non-image-containing questions, accuracy decreased from 72.94% (72.10–73.85) with complete clinical context to 53.53% (52.41–54.68) without the vignette and 37.36% (36.28–38.48) with answer options alone. OpenAI GPT-5.6 Terra achieved the highest full-context accuracy in both image-containing 81.80% (80.41–82.95) and non-image-containing 94.13% (93.30–94.97) questions. Removal of images alone did not significantly affect accuracy for any model, whereas removal of the clinical vignette significantly reduced accuracy for most models. When provided with only the answer options, without the question stem, clinical context, or images, the observed accuracy for every model exceeded the 25% expected from random selection. Specifically, the highest accuracy was 45.62% (44.24–47.24) for Claude Opus 4.8 on questions originally containing images and 46.09% (44.41–47.77) for GPT-5.6 Terra on questions without images. Although closed-source models achieved substantially higher accuracy than open-source models, all models performed above chance when provided answer options alone. BioMedBERT similarity remained high despite substantial differences in accuracy and was not significantly associated with accuracy by Spearman correlation (ρ=−0.17; P =0.29).
Conclusion
Clinical vignettes were the primary contributor to improving LLM performance on OITE whereas images had minimal effect. The high accuracy when only answer options were given suggests the models may be learning spurious correlations. Furthermore, BioMedBERT semantic similarity remained high across models despite wide differences in accuracy and did not consistently distinguish correct from incorrect responses. Thus, high examination performance may not directly reflect clinical reasoning, highlighting the importance of assessing how LLMs use the information provided when interpreting benchmark performance.