This retrospective preprint examines how different layers of clinical prediction — discrimination, probability estimates (calibration), and operating policy — transport when ICU delirium models are validated across databases. The authors performed a bidirectional external validation using models from five families trained and tested across two widely used critical care databases, MIMIC‑IV and eICU. The analysis emphasizes that deployment requires all three layers to transport, not only discrimination.
The study used de-identified retrospective data from MIMIC‑IV v3.1 and the eICU Collaborative Research Database v2.0, accessed through PhysioNet after credentialing, CITI training completion, and data use agreements. No identifiable patient information was accessed; no additional participant consent was required. The manuscript is a preprint and has not undergone peer review.
Five model families were evaluated in a bidirectional fashion between the two databases. The authors framed transportability as layered: (1) endpoint discrimination (ranking and AUROC), (2) transported probability estimates for individual predictions, and (3) an operating policy implemented via development-selected cutoffs and alerting behavior. Each layer was assessed separately to identify where transport breaks down.
Coarse-label internal AUROC for the evaluated models was reported in the range 0.87–0.92. When models were transferred in a source-only fashion to the external database (i.e., without retraining on target data), discrimination dropped, with external AUROC across models falling to 0.66–0.83. This decline demonstrates that strong internal discrimination does not guarantee maintained discrimination on external datasets.
The authors evaluated a repeated monitoring task conditioned on prior assessments (to detect persistence or recurrence). In this assessment-conditioned task, external AUROC improved substantially, reaching 0.76–0.94, indicating that incorporating assessment history can boost discrimination for monitoring tasks. However, when assessment history was removed from the feature set, external AUROC decreased by 0.16–0.32, showing a marked dependence on prior assessment features for transport in monitoring applications.
Expanding the feature set to include broader variables did not consistently improve transport performance across the evaluated models and tasks. The findings suggest that arbitrarily adding features is not a reliable strategy to improve external transport and that model dependence on specific features (such as assessment history) can dominate transport behavior.
Despite shifts in absolute probability estimates and policy performance, transported risk scores maintained useful ranking information. Specifically, the top predicted risk decile in transported scores concentrated future-positive ICU stays by 2.4–6.9-fold compared with lower-risk deciles. This indicates that ranking performance can persist across datasets even when probabilities are not well calibrated for the target site.
The authors evaluated practical operating policies derived from development-selected cutoffs. At the row-level (individual prediction rows), these cutoffs produced alerts on 0.3–2.0% of prediction rows, capturing 9.2–11.0% of future-positive rows. When alerts were deduplicated at the patient-stay level (to reflect clinically actionable alerts rather than repeated rows), 4.9–12.2% of stays were alerted and these alerted stays captured 43.9–49.4% of future-positive stays. These results show that alerting volume and the fraction of future-positive events captured depend strongly on how alerts are aggregated and on site-specific behavior.
The study demonstrates that transportability is layered: ranking (discrimination) can be more robust across databases, while probability estimates (calibration) and the operating policy (cutoffs and alert volumes) can remain site dependent. Assessment history can strongly influence monitoring performance, and expanding feature sets does not guarantee better transport. The authors argue that layered validation — separately assessing discrimination, probability transport, and policy behavior — is a prerequisite before prospective evaluation and deployment. They emphasize that layered validation alone is not proof of clinical benefit.
When considering external deployment of ICU delirium prediction models, rely on layered cross-database validation rather than internal AUROC alone. Expect potential declines in discrimination on external data, substantial sensitivity to features like assessment history, and shifts in alerting behavior that require site-specific calibration and policy adjustment. The preprint underscores the need for prospective evaluation to determine clinical benefit.