Multimodal Multiple Instance Learning for Fibroepithelial Tumor Diagnosis in Breast Ultrasound

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Fibroepithelial breast lesions include fibroadenomas (FAs) and phyllodes tumors (PTs), which are difficult to distinguish preoperatively due to their overlapping clinical and ultrasound features. While FAs are typically managed with observation, PTs require complete surgical excision with negative margins due to their risk of local recurrence and malignant progression. Because core needle biopsy can be inconclusive in about 15% of patients, some patients undergo unnecessary surgical procedures. Breast ultrasound (BUS) is a widely available, noninvasive imaging modality, but visual interpretation alone is limited by the overlap in lesion appearance. Artificial intelligence (AI)-based BUS offers the potential to improve the preoperative differentiation of PTs from FAs, reducing unnecessary operations while ensuring appropriate treatment for patients with PTs. Most deep learning approaches operate at the image level, whereas clinical diagnosis is made at the patient level by integrating multiple ultrasound views with clinical context. We propose a patient-level multimodal multiple instance learning (MIL) framework for fibroepithelial tumor diagnosis in breast ultrasound. Each patient is represented as a bag of BUS images encoded by an ultrasound foundation model (USFM) and aggregated using attention-based MIL to produce an image-derived patient score. Clinical variables, including age, lesion size, and echogenicity, are modeled using a gradient-boosted clinical learner. The image and clinical scores are then combined using logistic score fusion. In patient-level 5-fold cross-validation, logistic score fusion achieved 0.780 AUC-ROC for classification, improving over clinical feature-only gradient boosting (GB) by +0.033 (95% CI [+0.013, +0.053], p = 0.001), image-only MIL by +0.081 (95% CI [+0.042, +0.119], p < 0.001), and MLP-based end-to-end multimodal fusion by +0.035 (95% CI [+0.003, +0.067], p = 0.030). These results suggest that patient-level score fusion can combine complementary ultrasound and clinical evidence for PTs vs FAs.

Article activity feed