Automatic segmentation of medical images can speed clinical and research workflows, yet independent comparisons of available tools on common benchmarks are rare. This study benchmarks eight open-source automated thigh muscle MRI segmentation algorithms to address that gap. The authors focus on segmentation tasks that underpin volumetric analysis, radiomics, and fat fraction quantification—applications where segmentation errors can materially affect interpretation and downstream biomarker use. The work emphasizes independent evaluation on shared datasets to reveal genuine progress and practical limitations.
Eight open-source tools were evaluated. Six operate as independent segmentation methods, producing muscle labels directly from MRI volumes. Two additional algorithms are designed as auxiliary or refinement methods to be applied after muscles have been initially labeled by some other process. The study therefore includes a mix of end-to-end segmentation models and postprocessing/refinement tools to reflect approaches commonly used in practice.
Evaluation used three-dimensional MRI volumes drawn from three base datasets; the authors subsequently tested methods against derived datasets as well. The datasets included previously anonymized images, enabling the authors to perform the analysis without new institutional review board approval. Details on dataset provenance, links to the archived data, and code repositories are provided by the authors: related data and code are available on Zenodo and other archival services referenced in the manuscript.
The authors assessed algorithm performance using quantitative metrics appropriate to segmentation tasks and complemented these with qualitative usability assessment. Quantitative assessments measured segmentation accuracy across datasets; the manuscript reports that performance varied substantially between methods and across data sources. Qualitative usability considered practical aspects of deployment such as ease of use and likely integration into workflows, reflecting trade-offs beyond numeric accuracy alone.
Across the evaluated tools, the authors observed substantial variation in segmentation performance. Many methods exhibited reduced accuracy on pathological cases, indicating sensitivity to disease-related variation in muscle appearance. A clear pattern emerged favoring domain-specific models trained on large, task-focused datasets; these models consistently outperformed foundation models and newer general-purpose architectures on the thigh MRI segmentation tasks included in this benchmark. The two auxiliary/refinement algorithms were assessed in the context of their intended use after initial labeling, but the overall message is that method choice substantially affects results, particularly in nonnormal or underrepresented cases.
The findings highlight practical trade-offs that matter for both research and clinical adoption: accuracy, generalizability across datasets and pathologies, and ease of use. Because thigh muscle segmentation is often a precursor to volumetrics, radiomics, and fat fraction biomarkers in neuromuscular disease research and practice, segmentation errors can meaningfully influence downstream analyses and clinical interpretation. The authors therefore caution about deploying models that were not validated on representative pathological cases or underrepresented populations. The study suggests that domain-specific training with large, relevant datasets remains important for robust performance in these specialized tasks.
The authors made related data and code publicly available and archived them on Zenodo and other repositories; accession links are provided in the manuscript. Because the analyses were performed on previously anonymized open images, no new IRB approval was required, and the authors state that they followed relevant ethical guidelines. The preprint is released under a CC-BY 4.0 International license and the authors declared no competing interests.
This report is a preprint and has not undergone peer review; the authors explicitly state it should not be used to guide clinical practice. The manuscript documents that performance varied across datasets and that many methods showed reduced accuracy on pathological cases; specific numeric results, per-method rankings, or detailed metric values are reported in the full manuscript but are not repeated here. Readers interested in reproducing or extending the comparison are referred to the publicly archived data and code.
The study provides an independent, practical benchmark of open-source thigh muscle MRI segmentation tools, documenting strengths and weaknesses relevant to researchers and clinicians considering automated segmentation for volumetrics, radiomics, or fat fraction biomarker extraction. The principal conclusions are that (1) performance is heterogeneous across available tools and datasets, (2) domain-specific models trained on large, relevant datasets tend to outperform more general foundation models for this task, and (3) careful validation on pathological and underrepresented populations is essential before clinical deployment.