Multimodal Foundation Models for Medical Imaging - A Systematic Review and Implementation Guidelines

Shih-Cheng Huang
Malte Jensen
Serena Yeung-Levy
Matthew P. Lungren
Hoifung Poon
Akshay S Chaudhari

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Advancements in artificial intelligence (AI) offer promising solutions for enhancing clinical workflows and patient care, potentially revolutionizing healthcare delivery. However, the traditional paradigm of AI integration in healthcare is limited by models that rely on single input modalities during training and require extensive labeled data, failing to capture the multimodal nature of medical practice. Multimodal foundation models, particularly Large Vision Language Models (VLMs), have the potential to overcome these limitations by processing diverse data types and learning from large-scale unlabeled datasets or natural pairs of different modalities, thereby significantly contributing to the development of more robust and versatile AI systems in healthcare. In this review, we establish a unified terminology for multimodal foundation models for medical imaging applications and provide a systematic analysis of papers published between 2012 and 2024. In total, we screened 1,144 papers from medical and AI domains and extracted data from 97 included studies. Our comprehensive effort aggregates the collective knowledge of prior work, evaluates the current state of multimodal AI in healthcare, and delineates both prevailing limitations and potential growth areas. We provide implementation guidelines and actionable recommendations for various stakeholders, including model developers, clinicians, policymakers, and dataset curators.

Version published to 10.1101/2024.10.23.24316003 on medRxiv
Oct 23, 2024

A Survey of Contrastive Learning in Medical AI: Foundations, Biomedical Modalities, and Future Directions

This article has 6 authors:
1. George Obaido
2. Ibomoiye Domor Mienye
3. Kehinde Aruleba
4. Chidozie Williams Chukwu
5. Ebenezer Esenogho
6. Cameron Modisane
This article has no evaluationsLatest version Dec 26, 2025
Multimodal Machine Learning in Healthcare: A Tutorial and Review

This article has 4 authors:
1. Muntaqim Ahmed Raju
2. Priyanka Siddappa
3. Md Shifat Haider Al Amin
4. Ruizhe Ma
This article has no evaluationsLatest version Dec 16, 2025
QoQ-Med3: Robust Multimodal Clinical Analysis Foundation Model with Reasoning

This article has 10 authors:
1. David Dai
2. Jeannie She
3. Jiaee Cheong
4. Xing Han
5. Carl Harris
6. Haowen Wei
7. Farzan Vahedifard
8. Suchi Saria
9. Robert Stevens
10. Paul Liang
This article has no evaluationsLatest version Dec 30, 2025

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

A Survey of Contrastive Learning in Medical AI: Foundations, Biomedical Modalities, and Future Directions

Multimodal Machine Learning in Healthcare: A Tutorial and Review

QoQ-Med3: Robust Multimodal Clinical Analysis Foundation Model with Reasoning