Language models reflect clinical evidence but fail to adapt it to patients
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
Clinical language models must use evidence to produce the number and action required by a particular case. Whether numerical knowledge reliably becomes a correct clinical response is unclear.
Methods
NUMBERS evaluated 16 model configurations on 1,300 questions linked to public clinical evidence. Linked experiments tested prevalence updating, patient-specific estimates and clinical actions. The direct-action experiment compared cutoff recall plus action selection with a supplied complete rule across 50 rules, five models and 7,500 calls.
Results
Among diagnostic estimates outside the source-result tolerance, 77.1% remained within the evidence’s central 80% predictive range. Models updated prevalence-dependent quantities correctly in 78.5% of comparisons with inputs and a calculation request, versus 33.7% with clinical wording. Supplying inputs and requesting calculation raised patient-level near-target answers from 29.7% to 78.8% across 2,340 pairs. Models selected an incorrect action despite stating a cutoff that implied the correct action in 413 of 3,750 recall-arm calls (11.01%; 95% confidence interval, 8.93 to 13.17). Supplying the complete rule raised action accuracy from 85.63% to 99.41%, an improvement of 13.79 percentage points (95% confidence interval, 11.65 to 15.95).
Conclusions
Models often produced evidence-consistent numbers but failed to adapt them to a case or act consistently with their own stated cutoff. Explicit inputs, calculations and complete rules substantially improved performance in controlled prompts.
Funding
National Academy of Medicine, Agreement No. 2026A008797