This study evaluated the diagnostic performance of a research-grade spatial frequency domain imaging (SFDI) platform (Reflect RS) against a simplified commercial SFDI system (Clarifi RS) for early, quantitative assessment of burn wound severity. The objective was to quantify how reducing measurement dimensionality—fewer spatial frequencies and wavelengths, or acquisition features approximating a commercial device—affects the ability to classify burn severity when classifiers are trained using long-term healing outcomes.
Graded burns were created in a controlled porcine model and imaged 24 hours after injury using both SFDI platforms. The study leveraged the established capability of SFDI to detect optical property changes that correlate with tissue damage and healing potential. Regions of interest used for training and evaluation were defined by 28-day healing outcomes, providing a biologically relevant ground truth for supervised classification.
The primary dataset came from the research-grade Reflect RS system and was high-dimensional, incorporating multiple spatial frequencies and wavelengths. From this full Reflect dataset, the authors derived reduced-feature subsets to examine the effect of dimensionality reduction. In addition, datasets were constructed to mimic the acquisition features of the Clarifi RS commercial platform to allow direct comparison between a high-dimensional research system and a simplified commercial configuration.
Ground truth labels for each pixel were determined by mapping regions defined by healing outcomes at 28 days post-injury; pixel-level labels enabled fine-grained classifier training and performance assessment.
Pixel-level classifiers were trained on the labeled data and evaluated using a leave-one-subject-out cross-validation approach. This validation strategy measured generalizability by iteratively holding out data from one subject for testing while training on the remaining subjects. Performance was quantified using classification metrics with emphasis on the F1 score, enabling comparison across feature sets and simulated device configurations.
The highest classification performance was achieved using the full Reflect dataset, with mean F1 scores approaching 0.88. When the number of measured features was reduced, classification accuracy decreased relative to the full dataset. However, the reduction in performance was moderate: simplified models and datasets designed to mimic the Clarifi RS acquisition still produced robust results. For binary classification tasks, these reduced-feature and Clarifi-based datasets maintained F1 scores greater than 0.8.
These results indicate that while richer, high-dimensional SFDI measurements yield superior classification, clinically practical simplifications can retain acceptable diagnostic performance. The trade-off between measurement complexity and achievable accuracy was measurable but described as manageable in this controlled evaluation.
The retention of F1 > 0.8 in reduced-feature and Clarifi-like datasets supports the concept of developing application-specific, simplified SFDI systems for clinical burn assessment. Such systems could offer a balance between practical considerations (cost, acquisition speed, device complexity) and diagnostic utility, enabling broader clinical deployment while preserving useful predictive performance for healing outcomes.
This work is reported as a preprint and has not been peer reviewed; readers should interpret findings accordingly. The study used a controlled porcine model with pixel-level labels derived from 28-day healing, and the reported performance metrics reflect that experimental context. The authors disclosed funding from NIGMS (R01GM108634) and NIAMS (T32AR080622). Competing interests: Dr. Durkin is a co-founder of Modulim and is managed under institutional conflict-of-interest policies; the other authors declared no financial interests or commercial associations that represent conflicts relevant to the presented information.
(Details such as exact classifier types, hyperparameters, numerical breakdown of performance by class, sample sizes, and acquisition parameters were not reported in the source abstract and sections provided here.)