This study defines a standardized preclinical workflow and specific efficacy endpoints to evaluate candidate therapies for cutaneous neurofibromas (cNFs) using the Prss56Cre Nf1-KO mouse model. The model induces biallelic Nf1 loss together with tdTomato (Tom) reporter expression in Prss56-expressing Schwann cell (SC) lineages, enabling direct visualization of tumor SCs in vivo and in tissue sections.
Two experimental protocols were specified: a preventive protocol in which the investigational agent is delivered to younger mutant males prior to tumor onset, and a curative protocol in which treatment begins after mice present numerous mature cNFs. For both protocols macroscopic and IF endpoints are assessed at baseline and at treatment endpoint for inter-group (treated versus placebo) and intra-group (baseline versus endpoint) comparisons. For topical treatments, distinct skin areas are designated for drug versus vehicle application.
Macroscopic endpoints are assessed on 2D dorsal skin images acquired with a fluorescent magnifying loupe and LasX software. Defined macroscopic outcomes are: (1) number of cNFs, (2) total fluorescent surface area (TFSA) linked to Tom expression, and (3) fluorescence intensity. IF endpoints—measured on 14 μm confocal sections—are: (1) proportions of specific cell types (e.g., Tom+ SCs, fibroblasts, immune cells) and (2) proportions of area occupied by these markers within regions of interest.
A standardized image naming convention was instituted to support downstream automated analyses and metadata consolidation. Macroscopic filenames follow the pattern YYYY MM DD_drugname_beforeorafter_mousenumber; IF filenames follow YYYY MM DD_drug_mousenumber_cNFnumber_PANEL-DAPI-TOM. These conventions enable consistent CSV outputs from automated scripts and facilitate integration into a single metadata table for statistical workups.
Dedicated ImageJ/Fiji scripts were developed for both macroscopic and IF analyses to ensure reproducibility and to enable high-throughput processing. The macroscopic script quantifies Tom fluorescence across the dorsal skin to derive tumor counts, TFSA and intensity metrics. The IF script segments DAPI, Tom and antibody channels to estimate cell-type composition and area metrics according to the validated antibody panel used in this model.
Before batch processing, users adjust two main parameters in the IF script: contrast thresholds (to account for instrument- and acquisition-related variability) and the number/order of channels reflecting image acquisition sequence. The IF script is adaptable to the antibody panel reported by the authors.
The automated approach was validated against manual scoring by two blinded investigators using archived macroscopic and IF images from prior studies. Reproducibility of automated processing across computers was reported as perfect (ICC = 1). Inter-rater reproducibility for manual scoring was described as good for macroscopic outcomes and moderate for IF outcomes.
Correlation between automated and manual measurements was high: Spearman’s rho values reported were 0.92 for cNF number, 0.93 for TFSA, and 1.00 for fluorescence intensity; for IF outcomes, rho was 0.92 for Tom+ SC counts and 0.93 for Tom+ area. Bland–Altman analyses demonstrated no systematic bias between automated and manual methods; the largest mean difference was an underestimation of 9% for cNF number, with mean differences of 1% for TFSA, 0.5% for Tom+ area and 4% for Tom+ cell counts.
Automated scripts output standardized CSV tables with harmonized terminology compatible with the prescribed naming conventions. These outputs were consolidated into a single metadata table using RStudio for statistical analysis. The standardized metadata approach supports reproducibility and cross-study comparability by ensuring consistent variable naming and traceability from raw images to analyzed outcomes.
The workflow supports both preventive and curative study designs. In curative studies, topical or systemic compounds are administered after induction of cNFs (by cohort housing-induced skin microtrauma in this model); in preventive studies, treatment begins before tumor emergence. The same macroscopic and IF endpoints are applied across designs so that inter-study comparisons are feasible.
Automated analysis substantially reduced per-image processing time. Reported mean processing times showed large reductions versus manual scoring: macroscopic images that took minutes manually were processed in seconds automatically; IF images reduced from roughly 9 minutes manually to near-instant automated times per image. Statistical analyses used ICC for reproducibility, Spearman correlation for concordance with manual scoring, and Bland–Altman plots for bias assessment. All data supporting these results are publicly available via the authors’ Synapse repository.
The authors note that the Prss56Cre Nf1-KO model produces flat, punctate cNFs that are not amenable to ruler- or caliper-based measures, supporting the need for fluorescence-based quantification. While the automated workflow addresses objectivity and throughput, manual review and adjustment of contrast thresholds remain necessary to account for acquisition variability. The group is evaluating in vivo benchtop fluorescence systems for additional live imaging options. Overall, the validated ImageJ-based pipeline provides a reproducible, time-efficient framework to quantify cNF burden and composition in this mouse model, facilitating more rigorous preclinical efficacy testing and cross-study comparisons.