This study evaluates how different segmentation strategies—specifically conservative (eroded) versus extensive (dilated) tumour boundaries—affect the stability of first-order radiomic features derived from ADC (apparent diffusion coefficient) maps in paediatric brain tumours. The central aim was to quantify feature variability introduced by typical annotation biases and to determine the downstream effects on machine-learning diagnostic models.
Researchers used a retrospective cohort drawn from the Imaging of Tumours (IoT) study, with MRI performed at Birmingham Children’s Hospital as part of standard-of-care imaging prior to surgical intervention. From an initial 131 participants with pre-surgical DWI, exclusions reduced the analysis set to 106 participants. Exclusions included images with significant artefacts (N = 13), diffuse tumours with unclear margins (N = 9), and three participants removed after ROI drawing due to insufficient tumour volume (< 100 voxels in ≥2 axial slices).
The final cohort encompassed 25 diagnostic classes. The three largest diagnostic groups were pilocytic astrocytomas (N = 30), medulloblastomas (N = 27), and ependymomas (N = 12). Diagnoses were assigned by WHO histology where available or by radiological diagnosis when biopsy was not performed.
Ground-truth (GT) tumour ROIs were used as the baseline segmentation. To simulate common annotation biases, GT ROIs were systematically eroded to represent conservative boundary-drawing and dilated to represent extensive boundary-drawing. Nineteen first-order radiomic features were extracted from ADC maps within each ROI variant to characterise intensity-based properties of the tumour region.
Feature stability under erosion and dilation was quantified and compared. For 18 of the 19 first-order features, dilation produced significantly greater variability than erosion (p < 0.01). Effect size analysis showed a large effect (d > 0.8) for 11 features, indicating a substantial impact of including ambiguous boundary regions on many first-order ADC statistics. The degree of sensitivity to ROI perturbation varied by diagnosis, with the strongest relationship observed in pilocytic astrocytomas.
To evaluate practical consequences, extracted features from GT and perturbed ROIs were used to train diagnostic machine-learning models. Models trained using features from GT ROIs suffered reductions in classification accuracy when tested on perturbed ROIs: accuracy decreased by 3.8 ± 0.8% for low-level erosion and by 5.6 ± 0.9% for low-level dilation. These results demonstrate that segmentation strategy can materially alter model performance, and that over-segmentation (including more ambiguous boundary voxels) tends to have a larger negative effect than under-segmentation.
The authors tested strategies to improve robustness to variable ROI drawing. Augmenting the training dataset with features derived from eroded and dilated ROIs, combined with a feature-selection step that prioritised stable features, partially mitigated the adverse effects of segmentation variability. With this combined approach, model accuracy drops relative to GT-trained models were reduced to 1.4 ± 0.7% for erosion and 2.9 ± 0.3% for dilation. This indicates that training-time exposure to annotation variability and pruning to stable features improves generalisability of diagnostic models.
Key conclusions from the study are:
Including ambiguous boundary voxels via dilation introduces greater instability in first-order ADC radiomic features than excluding them via erosion.
Conservative segmentation approaches (excluding uncertain boundary regions) generally produce less variability in extracted features and therefore may be preferable when defining ROIs for radiomic analyses intended for clinical translation.
Diagnostic models are sensitive to segmentation variability, but model robustness can be enhanced through training augmentation with perturbed ROIs and by selecting features that remain stable across segmentation perturbations.
These findings provide practical guidance for radiomic workflows in paediatric brain tumour imaging, suggesting that careful consideration of segmentation strategy and explicit handling of annotation variability should be incorporated during model development and validation.
The study used child patient data; as such, full raw datasets cannot be shared publicly and access is restricted under the study’s ethical approvals. The authors made code and extensive result tables available in a public GitHub repository referenced in the paper, which include image-processing and machine-learning/statistical analysis scripts plus raw values used to build figures and calculate summary statistics. Contact details for data access inquiries are provided in the original article. The study was performed with informed consent and appropriate NHS research ethics approval.
While this work quantifies first-order feature sensitivity and demonstrates mitigation strategies for model robustness, application to higher-order radiomic features, different MRI contrasts, and multi-centre datasets would be needed to generalise recommendations. The authors note the importance of formalising segmentation guidelines and incorporating annotation variability into model development to improve clinical applicability of radiomic diagnostics.