Critically ill patients with cancer who develop sepsis have high in-hospital mortality, yet externally validated prognostic models specifically for this group are limited. The investigators aimed to build an interpretable machine learning framework that uses routinely available first-day intensive care unit data to predict in-hospital death. The work focused on a predetermined 24-hour landmark after ICU admission and excluded diagnosis-derived variables from the same admission to avoid leakage.
The development cohort was derived from MIMIC-IV version 3.1 and included ICU stays meeting inclusion criteria: adults with cancer, ICU length of stay ≥ 24 hours, and meeting a prespecified operational definition of sepsis. The 24-hour post-admission timepoint served as the prediction landmark.
Fourteen predictive models plus a dummy baseline model were trained and compared. A total of 345 predictors drawn from first-day ICU data were retained after exclusion of same-admission diagnosis-derived features. Model training and internal validation were performed within MIMIC-IV; frozen pipelines and thresholds derived from training were then applied without refitting or recalibration to an external dataset (eICU-CRD v2.0) for locked validation.
The MIMIC-IV development cohort comprised 3,729 ICU stays, split into a training set (n = 2,983) and an internal validation set (n = 746). External validation used the eICU-CRD dataset with 611 eligible stays. Both datasets are de-identified, credentialed-access resources available through PhysioNet.
Predictor selection prioritized variables available within the first 24 hours of ICU admission. Same-admission diagnosis-derived variables were removed to reduce target leakage. From the remaining candidate features, 345 predictors were retained for model building. The training process produced frozen preprocessing and modeling pipelines as well as decision thresholds that were preserved for external locked testing.
Fourteen modeling approaches were evaluated; based on internal performance and interpretability assessments, gradient boosting was selected as the primary model for reporting. Other models, including random forest, were retained for secondary comparisons in external testing.
In internal validation within MIMIC-IV, the selected gradient boosting model demonstrated strong discrimination and calibration metrics: AUROC 0.8480 (95% CI, 0.8173–0.8755), area under the precision-recall curve (AUPRC) 0.6984, and a Brier score of 0.1346. These metrics indicate useful internal discrimination based on first-day ICU data in this development cohort.
Locked external validation applied the frozen MIMIC-IV-derived pipelines and thresholds directly to the eICU-CRD cohort without refitting or recalibration. Under this strict external test, performance declined: gradient boosting achieved AUROC 0.7483 (95% CI, 0.7026–0.7919), AUPRC 0.6243, and Brier score 0.1709. In secondary comparisons, a random forest model showed the highest external AUROC (0.7731) among the approaches evaluated.
The authors interpret the drop in performance as evidence that models trained on single-center or single-source data can lose discrimination and calibration when transported to multicenter, external settings without adaptation.
Model explainability analyses identified several consistently important first-day predictors. These included components of the Glasgow Coma Scale, body temperature, lactate dehydrogenase, patient age, respiratory rate, oxygen saturation, blood urea nitrogen, and serum lactate. The emphasis on these physiologic and laboratory variables supports clinical face validity for mortality risk stratification in septic cancer patients.
Strengths reported by the authors include use of large, credentialed ICU databases (MIMIC-IV and eICU-CRD), a clear 24-hour prediction landmark, exclusion of same-admission diagnosis-derived features to limit data leakage, and a locked external validation strategy applying frozen pipelines and thresholds.
Limitations evident from the reported results include reduced discrimination and calibration under locked external validation, indicating limited immediate transportability. The report notes that multicenter validation, recalibration, threshold assessment, and prospective evaluation are required before any clinical deployment. The manuscript is a preprint and has not undergone peer review; the authors caution that findings should not guide clinical practice at this stage.
The study used de-identified records from MIMIC-IV v3.1 and eICU-CRD v2.0 accessed through PhysioNet after required credentialing and training. Because the work was secondary analysis of de-identified datasets, no additional institutional review board approval or individual informed consent was required. Patient-level records and derived extracts cannot be redistributed by the authors because the datasets are third-party controlled-access resources; however, credentialed researchers can request access through PhysioNet.
First-day ICU data enabled an interpretable machine learning model with good internal discrimination for predicting in-hospital mortality among ICU patients with cancer and sepsis. Performance decreased under a locked multicenter external validation strategy, highlighting the need for further work before clinical implementation. The authors recommend multicenter validation, recalibration of model outputs, assessment of operating thresholds, and prospective evaluation to establish generalizability, calibration, and clinical utility.