Benchmarking Open-Source Vision-Language Models for Brain Metastasis Assessment on Single-Slice Contrast-Enhanced MRI
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Purpose
Open-source vision-language models (VLMs) can be locally deployed without external internet access, potentially enhancing data security. This study compared the diagnostic performance of general-purpose and medical-purpose open-source VLMs and evaluated their ability to characterize brain metastases on contrast-enhanced (CE) MRI.
Materials and Methods
Sixty lesion-positive axial CE T1-weighted images and sixty matched lesion-negative images from 60 patients were analyzed using three general-purpose VLMs–InternVL3-8B, Qwen2.5-VL-7B-Instruct, and MiniCPM-V-4.5–and three medical-purpose VLMs–MedGemma-4B-it, LLaVA-Med v1.5, and HuatuoGPT-Vision-7B. Lesion detection performance was assessed using sensitivity, specificity, and balanced accuracy. On lesion-positive images, accuracy was evaluated for lesion count, laterality, anatomic location, enhancement pattern, necrosis, vasogenic edema, and mass effect. Model differences were assessed using Cochran’s Q tests followed by pairwise McNemar tests with Benjamini-Hochberg correction.
Results
The median age of the study patients was 67 years (IQR, 61.0–70.5 years), and 35 patients were male (58.3%). MiniCPM-V-4.5 showed the most balanced diagnostic performance, with a sensitivity of 78.3% (95% CI, 66.4–86.9%) and a specificity of 85.0% (95% CI, 73.9–91.9%), and significantly higher balanced accuracy than all other models. Significant overall differences were observed for lesion count, laterality, location, enhancement pattern, necrosis, and mass effect, but not for vasogenic edema (FDR-adjusted P = 0.056). HuatuoGPT-Vision-7B and MedGemma-4B-it showed relatively consistent accuracy across multiple image assessment tasks, although their performance remained modest.
Conclusion
Our study demonstrated substantial heterogeneity in the performance of open-source VLMs in brain metastasis evaluation, and medical-purpose VLMs did not outperform general-purpose VLMs.