Early identification of patients at high risk of in-hospital mortality after intensive care unit (ICU) admission can guide clinical decision-making and resource allocation. The investigators set out to develop a multimodal artificial intelligence deep learning model that integrates structured and unstructured clinical data to predict subsequent inpatient mortality using information available during the first 24 hours of ICU admission.
The model development used large, multicenter critical care datasets. Training and internal development occurred using the MIMIC datasets. Key performance measures reported were area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), and Brier scores. Statistical comparison of AUROCs used the DeLong test where applicable.
Four datasets were used: MIMIC-III, MIMIC-IV, eICU, and HiRID. Across these sources the study included a total of 203,434 ICU admissions drawn from more than 200 hospitals with data spanning 2001 to 2022. Observed in-hospital mortality rates across the datasets ranged from 5.2% to 7.9%.
The multimodal model combined multiple input types:
The model was developed on the MIMIC datasets and designed to predict risk of subsequent inpatient mortality after the initial 24-hour ICU period. Development included integrating the distinct modalities into a single predictive architecture; specific modeling architecture details and hyperparameters were not reported in the abstract.
External validation was performed in multiple ways. A temporally separated MIMIC population was used for validation, and two additional external datasets were applied: HiRID and eICU. Within the eICU dataset, validation was reported across eight different institutions to assess site-level generalizability. The abstract reports aggregated performance metrics and the range of AUROCs across these external sites.
For the model integrating structured data points, reported performance on relevant test data included an AUROC of 0.92 (95% confidence interval [CI] 0.90–0.93), an AUPRC of 0.53 (95% CI 0.49–0.57), and a Brier score of 0.19 (95% CI 0.18–0.20).
External validation within eight institutions of the eICU dataset produced AUROCs ranging from 0.84 to 0.92, indicating the model maintained discrimination across different sites.
In a subgroup analysis restricted to patients with available clinical notes and imaging, adding those modalities to structured inputs improved discrimination and calibration. Measured changes were:
These results indicate measurable benefit of incorporating unstructured notes and imaging into the mortality prediction model for the subset of patients with those data available.
The study demonstrates that a deep learning model combining structured time-invariant and time-variant data with unstructured clinical notes and chest X-rays can predict in-hospital mortality after the first 24 hours of ICU admission. Performance was strong for structured-data models and remained robust across external validation sites, with further improvement when notes and imaging were included for eligible patients. The authors highlight two central points: the value of integrating multiple patient information sources for prediction, and the importance of external validation across institutions to assess generalizability.
Clinical implementation considerations, operational requirements, and detailed model architecture or prospective impact on care processes were not described in the abstract.
The abstract reports that Dr. Gabriel’s institution, the University of California, has received funding and/or product for research from several organizations. No further conflict of interest details were provided in the abstract.