Basal cell carcinoma (BCC) management follows a staged decision sequence that includes triage, pathological subtyping, and depth assessment. At each step different information is available, and decision needs differ. The authors developed an endpoint-aligned artificial intelligence (AI) framework intended to match non-invasive inputs to the specific decision points in this clinical sequence, with the goal of supporting biopsy-sparing assessment and personalized treatment planning.
The study used internal and external patient cohorts: 1,459 internal patients and 995 external patients were included for development and evaluation. A geographically distinct cohort was also evaluated and reported separately. Individual-level clinical and dermoscopic photographs from the XJTU and Baoji cohorts are not publicly available because of patient confidentiality. The PAD-UFES-20 dataset is publicly available and was referenced. Source data underlying figures and tables (derived numerical values without patient-level images) can be requested from the corresponding author.
Ethical approvals were obtained from The Second Affiliated Hospital of Xi'an Jiaotong University (approval number 2025-215) and Baoji Central Hospital (approval number BZYL2026-65). The authors declared that patient consent and oversight procedures were followed.
The AI framework was explicitly designed to support multiple endpoints that mirror real-world clinical decisions for suspected BCC. Non-invasive inputs (clinical and dermoscopic images and other unspecified modalities as used in the study) were paired with endpoint-specific models or configurations so that the information available at each step informed the corresponding prediction task. The endpoints reported include triage (determining whether a lesion is likely BCC and requires further assessment), risk stratification, and tumor thickness estimation relevant to treatment planning.
Triage performance, reported as macro-AUROC, was very high in closed-set evaluations: 0.995 on the internal cohort and 0.978 on the external cohort. In a geographically distinct cohort triage macro-AUROC decreased to 0.853, reflecting domain shift. Risk-stratification AUROCs were reported as 0.943 and 0.899 for the evaluated settings.
These results highlight strong closed-set discrimination for triage and risk stratification but also show performance degradation in a geographically separate population, underscoring the importance of scope control and external validation.
For the endpoint of tumor thickness estimation, the authors evaluated different input modality configurations. Surprisingly, a configuration using dermoscopy images alone produced higher precision than a configuration using all available modalities: precision of 0.949 with dermoscopy alone versus 0.881 when all modalities were included. This finding suggests that modality selection and endpoint-specific input design can materially affect predictive precision for biopsy-sparing tasks.
The AI framework was benchmarked against human experts. On matched cases, AI performance exceeded the mean performance of 19 dermatologists for every prespecified primary metric; the reported P values for these comparisons were all ≤ 0.014. The source text does not provide further detail here about the exact metrics compared per endpoint or case selection criteria beyond stating a matched-case comparison and the statistical significance bounds.
The authors examined the effect of local adaptation (recalibration or local retraining) on model behavior. Local adaptation increased in-scope accuracy substantially—from 0.790 to 0.954. However, this adjustment changed how the system handled out-of-scope inputs: it shifted routing toward "no-further-assessment" classes for inputs outside the intended scope, and overall sensitivity for those cases fell from 0.953 to 0.697. This trade-off illustrates that improving closed-set accuracy with local adaptation can worsen safety-relevant behavior for out-of-scope data unless scope-control mechanisms are also addressed.
To address out-of-scope handling, the study implemented a validation-locked Mahalanobis gate—a statistical mechanism to screen inputs against the distribution of training/validation data. Applying this gate enriched sensitivity among the accepted (in-distribution) cases to 0.775 at 0.791 coverage. However, the gate only partially mitigated residual out-of-scope routing errors; it improved sensitivity among accepted cases but did not fully prevent erroneous routing of out-of-scope inputs.
The authors emphasize a separation between closed-set performance and scope-control performance: excellent performance on in-distribution test sets does not guarantee safe routing or reliable behavior when the model encounters out-of-scope inputs or data from different geographic settings. The findings support endpoint-specific validation when developing biopsy-sparing AI tools for BCC diagnosis and personalized treatment planning.
Limitations reported or implied by the source text include restricted availability of individual-level images due to confidentiality and observed performance drops in geographically distinct cohorts; the source does not include full methodological details, exact model architectures, or exhaustive metrics for every endpoint in this summary abstract. The study is a preprint and has not been peer reviewed; the authors explicitly state it should not be used to guide clinical practice without further validation.
An endpoint-aligned AI approach using non-invasive inputs demonstrated strong closed-set discrimination for triage, risk stratification, and thickness estimation in suspected BCC and outperformed a dermatologist mean on matched cases. However, local adaptation and scope-control trade-offs highlight the need for careful external validation, explicit out-of-distribution detection, and endpoint-specific performance criteria before clinical deployment.