Accuracy Overstates Evidence Grounding and Abstention Reliability in Mammography Vision-Language Models
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Answer accuracy alone cannot determine whether a Vision-Language Model (VLM) relies on clinically relevant mammographic evidence or recognizes when that evidence is unavailable. We introduce an evidence-grounded selective evaluation benchmark that evaluates Pathology classification and Abnormality identification together with label-aware lesion localization. Each original image–task instance is paired with a lesion-removal counterfactual and, when feasible, a size-matched non-lesion random-removal control. Lesion evidence is removed by local tissue reconstruction, while the random-removal view provides a task-irrelevant regional perturbation baseline of comparable spatial extent. We evaluate 16 general-purpose and medically specialized VLMs, including open-source and proprietary models, under a common protocol with the same prompt, image conditions, label spaces, and structured-output requirements. The results show substantial gaps between classification performance and evidence-grounded reliability. InternVL3.5-8B achieved the highest Pathology Macro-F1 at 53.75%, Huatuo Vision-34B achieved the highest Abnormality Exact Match at 52.44%, and GPT-5.6 Sol achieved the highest Grounded Answer Success at 21.99%. Lingshu-7B achieved the highest Counterfactual Specificity at 35.19%, whereas GPT-5.6 Sol achieved the highest Joint Reliability at 6.47%. These findings suggest that answer accuracy and abstention behavior can substantially overstate the reliability of current VLMs when predictions are not verified against the visual evidence on which they should depend.