Robustness of Deep Learning Segmentation Across Heterogeneous Uterine MRI Datasets: A Multi-Task Comparison of U-Net, Swin-UNETR, MedNeXt, and nnU-Net
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Reliable medical image segmentation requires performance that is stable across tasks, acquisition settings, and evaluation domains. We evaluated U-Net, Swin-UNETR, and MedNeXt as manually configured architectures and nnU-Net as a self-configuring reference across three T2-weighted uterine MRI datasets: public multiclass uterine anatomy and fibroid segmentation (UMD, n=300), institutional endometrial cancer tumor segmentation (n=206), and institutional uterine mass lesion segmentation (n=234). A separate UMD-style cohort (n=12), relabeled to the same annotation ontology, was used to assess external domain shift. Fixed data partitions, fold ensembling, and matched evaluation were used across models. Performance was assessed with Dice, 95th-percentile Hausdorff distance, average symmetric surface distance, absolute volume difference, and paired bootstrap comparisons with within-dataset Holm correction. MedNeXt was the strongest manually configured architecture, whereas nnU-Net achieved the highest Dice on all internal datasets: 0.761 for UMD, 0.746 for endometrial cancer, and 0.814 for uterine mass segmentation, compared with 0.722, 0.726, and 0.789 for MedNeXt. The nnU-Net-MedNeXt separation was largest for multiclass UMD segmentation and smaller for the binary tasks. All models degraded substantially on external testing; Dice was 0.542 for nnU-Net, 0.490 for MedNeXt, 0.396 for U-Net, and 0.287 for Swin-UNETR. These findings show that architecture, task definition, and dataset-adaptive pipeline configuration jointly influence uterine MRI segmentation, while external domain shift remains a major source of failure.