Benchmarking open-source automated thigh muscle MRI segmentation algorithms

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Automatic segmentation of medical images could accelerate clinical and research workflows, yet existing tools are rarely compared independently on shared benchmarks, making it difficult to assess genuine progress. This is particularly consequential for limb muscle segmentation in magnetic resonance imaging (MRI), which underpins volumetric analysis, radiomics, and fat fraction quantification used as biomarkers in neuromuscular disease – applications where segmentation errors can directly affect interpretation. Despite the availability of multiple open-source segmentation tools spanning diverse architectures, no independent head-to-head comparison has appeared in the literature. We address this gap by evaluating eight tools, six of which run independently, and the other two of which are meant to be auxiliary algorithms once muscles have been somehow initially labeled. We evaluated these tools on three-dimensional MRI volumes across three base datasets and then later against derived datasets, assessing both quantitative metrics and qualitative usability. Our results reveal substantial performance variation across methods and datasets, with many methods showing reduced accuracy on pathological cases. Domain-specific models trained on large datasets consistently outperformed foundation models and newer general-purpose architectures. Our results reveal practical trade-offs between accuracy, generalizability, and ease of use that are critical for clinical adoption, while raising concerns for deployment on underrepresented populations and pathologies.

Article activity feed