Magnetic resonance imaging (MRI) detection and quantification of white matter lesions (WML) are central to diagnosing and monitoring multiple sclerosis (MS). Portable ultra‑low field (pULF) MRI operating at 64 millitesla (mT) has demonstrated capacity to visualize WML with at least one dimension greater than 4 mm. An automated segmentation tool designed for pULF‑MRI could provide standardized, objective WML volumetrics to support clinical assessment and research.
This study used same‑day paired pULF (64 mT) and high‑field (HF, 3T) MRI acquired from 84 adults with MS or suspected MS. The sample had a mean age of 48 ± 13 years and included 62 females. Image contrasts acquired included T2‑FLAIR and T1‑weighted (T1w) sequences at both field strengths when available.
Reference WML masks on pULF T2‑FLAIR images were manually annotated for all scans. These pULF annotations were confirmed using registered HF T2‑FLAIR images. Separate HF reference WML segmentations were also generated to support method training and evaluation.
Six automated approaches were applied to pULF scans:
Two models were trained within the nnU‑Net framework: one using T2‑FLAIR only (nnU‑Net‑FL) and one using both T1w and T2‑FLAIR (nnU‑Net‑FL/T1). The same configuration was applied to PLAn to produce PLAn‑FL and PLAn‑FL/T1. Together with MIMoSA and WMH‑SynthSeg, these produced six automated segmentation outputs for comparison.
Automated segmentation outputs were compared to the manual pULF reference masks using Dice Similarity Coefficient (DSC) as the primary overlap metric. Associations between estimated WML volumes and clinical disability measures—the Expanded Disability Status Scale (EDSS) and the Scripps Neurologic Rating Scale (SNRS)—were also assessed. Age‑adjusted analyses examined whether automated WML volume estimates related to clinical measures after accounting for age.
Mean DSC values (mean ± SD) comparing automated outputs to pULF reference masks were reported as follows:
PLAn‑FL statistically outperformed MIMoSA, WMH‑SynthSeg, nnU‑Net‑FL, and nnU‑Net‑FL/T1 with reported p‑values indicating significance (for example, p < 0.0001 for comparisons versus MIMoSA and WMH‑SynthSeg). The two nnU‑Net variants performed similarly to each other and were intermediate between PLAn‑FL and the lower‑performing methods.
Higher WML volumes measured in both the pULF and HF reference masks were correlated with worse clinical status on EDSS and SNRS. Automated WML volumes derived from WMH‑SynthSeg, nnU‑Net‑FL, nnU‑Net‑FL/T1, PLAn‑FL, and PLAn‑FL/T1 also correlated with EDSS and SNRS. MIMoSA‑derived WML volumes did not show this association.
After adjusting for age, WML volumes estimated by WMH‑SynthSeg, the nnU‑Net variants, and the PLAn variants retained significant associations with EDSS and SNRS. These results indicate that several automated methods produced WML burden estimates that reflected clinical and radiological disease severity.
In this cohort, deep‑learning approaches—particularly nnU‑Net and PLAn—performed best for segmenting WML on 64 mT pULF‑MRI, providing more accurate quantitative estimates of WML burden than the evaluated machine‑learning method. The correspondence between algorithm‑derived WML volumes and measures of disability (EDSS and SNRS) supports the potential clinical relevance of automated quantification from pULF images.
Given the mobility and lower cost of pULF‑MRI, coupling it with effective automated segmentation could expand access to imaging in clinical trials and routine practice, especially for participants facing travel, mobility, or logistic barriers. The authors highlight this potential as a rationale for further development and application of pULF imaging with tailored deep‑learning segmentation.
The work received ethical approval from the Institutional Review Board of the National Institutes of Health. Funding sources included the Intramural Research Program of NINDS, the NIHR Oxford Health Biomedical Research Centre, and the National Multiple Sclerosis Society. The manuscript is posted as a preprint on medRxiv and has not been peer reviewed; the authors note it should not be used to guide clinical practice without further validation.