Patients with EGFR-mutant lung adenocarcinoma who develop resistance to first-generation tyrosine kinase inhibitors (TKIs) without acquiring the T790M mutation have limited standard systemic options. Immune-checkpoint inhibitors combined with chemotherapy (ICI-chemotherapy) are used in this setting but demonstrate variable and often modest efficacy. Existing biomarkers such as PD-L1 have limited predictive value and obtaining post-resistance tissue for advanced profiling is often impractical. The investigators hypothesized that routine post-resistance, pre-immunotherapy CT scans may contain latent imaging signatures reflecting tumor biology and immune microenvironment changes that predict benefit from ICI-chemotherapy. The primary objective was to develop and externally validate end-to-end deep learning models using CT images to predict objective response after ICI-chemotherapy and to assess PFS stratification based on model output.
The study screened 750 patients across three centers and, after applying predefined inclusion and exclusion criteria, produced a final analyzable cohort of 490 patients with complete CT imaging and outcome labels. Exclusion reasons included missing follow-up or treatment response labels (n = 32), poor-quality or incomplete CT imaging (n = 74), and ineligibility per predefined criteria such as absence of sensitizing EGFR mutation, presence of T790M or other driver mutations, prior systemic therapy before EGFR-TKI initiation, or incomplete clinical records (n = 154). All included patients underwent post-resistance molecular profiling to confirm absence of the EGFR T790M mutation and exclusion of other driver alterations; tissue-based NGS was the primary method, with circulating tumor DNA used when tissue was unavailable.
The final cohort was split into a training set (n = 326) for model development and two institutionally independent external validation cohorts: validation cohort 1 (n = 70) and validation cohort 2 (n = 94). Neither validation cohort contributed to model training, hyperparameter tuning, threshold selection, or model selection.
Comprehensive clinical and pathological variables were retrospectively collected, including age, sex, smoking history, ECOG performance status, clinical stage, tumor location, specific EGFR mutation subtype (exon 19 deletion or L858R), treatment regimen, and subsequent therapies where available. PD-L1 tumor proportion score was recorded when available but testing was not uniform across centers; PD-L1 analyses were treated as exploratory. Pretreatment evaluation included ECOG status, contrast-enhanced CT of chest/abdomen/pelvis within 4 weeks before treatment, obligatory brain imaging, and baseline laboratory testing to ensure fitness for ICI-chemotherapy.
Tumor segmentation was performed independently by two observers blinded to outcomes. Interobserver reproducibility assessed on a 30-case subset yielded ICC values ranging from 0.973 to 1.000, with a median ICC of 0.998, supporting excellent segmentation consistency. Disagreements were resolved by consensus. ROI extraction procedures differed by model architecture as described below.
Three architectures based on the ResNet-101 backbone were developed and compared:
2D ResNet-101: For each orthogonal view (axial, sagittal, coronal), the single tumor-centered slice with the largest tumor area was cropped to a 256 × 256 patch and used as input. Three view-specific 2D models were trained, feature representations fused, and a final classifier produced the predicted probability of response.
2.5D ResNet-101: For each view, five adjacent slices (central slice with two neighboring slices on each side at 1-mm intervals) were stacked as five input channels to provide limited through-plane spatial context while retaining 2D pretrained-weight advantages. ROIs were defined as bounding rectangles across the five slices, resized to 256 × 256, and processed using a modified first convolutional layer accommodating five channels. Three view-specific 2.5D models were trained and fused.
3D ResNet-101: Tumor ROIs were enclosed in bounding cubes resampled to 96 × 96 × 96 voxels and processed volumetrically. The 3D model was trained from scratch because ImageNet-pretrained weights are not directly transferable to standard 3D medical networks.
Standard data augmentation included random flipping and center cropping. ImageNet initialization was applied to 2D and 2.5D backbones; the 3D model used random initialization.
All model training, hyperparameter tuning, and threshold derivation were restricted to the training cohort. The response-positive group (CR or PR by RECIST 1.1) and response-negative group (SD or PD) in the training cohort numbered 118 and 208 patients, respectively, indicating moderate class imbalance; no oversampling, undersampling, or class weighting was applied. The Youden index in the training cohort was used to derive classification thresholds, which were then locked and applied unchanged to both validation cohorts.
Statistical analyses included receiver operating characteristic area under the curve (AUC), accuracy, sensitivity, specificity, PPV, NPV, and F1 score. Progression-free survival (PFS) was defined as time from ICI-chemotherapy initiation to radiological progression or death and was used for exploratory risk stratification based on model score. Overall survival analyses were not reported due to potential confounding by subsequent therapies. The investigators noted that incomplete harmonization of some clinical variables across centers precluded construction of formal clinical-only or combined clinical-imaging models in this manuscript.
The model training endpoint was objective response rate (ORR) defined by RECIST 1.1 as best overall response (CR/PR vs SD/PD). Labels were binary with 1 indicating response-positive disease (CR or PR) and 0 indicating response-negative disease (SD or PD). PFS served as the time-to-event endpoint for survival stratification analyses derived from model outputs.
Across the entire 490-patient cohort, the observed ORR was 34.1%. For the 2.5D axial model specifically, reported AUCs using the locked source prediction files were 0.885 in the training cohort, 0.819 in validation cohort 1, and 0.863 in validation cohort 2. These results indicate promising discriminatory performance for predicting objective response to ICI-chemotherapy in the post-TKI resistance setting. The study also evaluated the model score for PFS stratification, though specific PFS hazard ratios or Kaplan–Meier statistics were not reported in the provided source text. The investigators emphasized that thresholds were not re-optimized in validation cohorts and that model outputs were evaluated both as classification probabilities and for time-to-event stratification.
The authors acknowledge important limitations: the retrospective study design, incomplete availability and limited harmonization of some clinical variables that prevented building clinical-only or combined models, and absence of prospective validation or clinical-variable benchmarking in this report. While a CT-based 2.5D deep learning model demonstrated promising predictive performance for ICI-chemotherapy response after EGFR-TKI resistance, the manuscript concludes that prospective validation and direct comparison with clinical-variable models are necessary before clinical implementation.