This preprint examines the potential of offline reinforcement learning (RL) to produce and evaluate treatment-support policies for ICU sepsis management using previously logged clinical data. Offline RL is attractive in critical care because it enables policy learning without prospective exploration, which would be unsafe for patients. The central question addressed is whether offline RL policies maintain a stable, action-sensitive decision-support signal when evaluated on increasingly severe, out-of-distribution (OOD) patient cohorts.
The analysis uses the MIMIC-III Clinical Database (version 1.4) as the source of logged ICU trajectories. MIMIC-III is a credentialed-access dataset available through PhysioNet, and the authors note they are not permitted to redistribute the underlying patient-level data. To probe robustness to severity shifts, the authors constructed three severity-enriched OOD test mixtures from the MIMIC-III benchmark dataset. The manuscript reports results separately for each of these mixtures to evaluate how offline RL methods behave as the proportion of severe patients in test mixtures increases.
The study evaluates standard offline RL methods and applies a shared learned-dynamics off-policy evaluation (OPE) protocol. Under this protocol, dynamics are learned from logged data and used to generate model-based evaluations of candidate policies. The learned-dynamics OPE framework serves as an OOD stress test to determine whether offline policies continue to produce a consistent decision-support signal when applied to cohorts with different severity mixes than those seen in the training data.
As the fraction of severe-OOD patients increases across the three test mixtures, the observed clinical terminal survival in the logged data declines. The manuscript reports observed survival values of 67% for the mixture with the lower severe-OOD ratio and 49% for the mixture with the higher severe-OOD ratio. These observed survival rates reflect the clinical outcomes actually recorded in the MIMIC-III trajectories used to form each severity-enriched test mixture.
For each severity-enriched mixture, the best-performing offline RL method—under the learned-dynamics OPE protocol—received substantially higher model-predicted terminal survival values than the observed clinical survival. Reported model-predicted terminal survival values for the best offline method were 87%, 86%, and 85% across the three mixtures, respectively. The authors emphasize that observed clinical survival and model-predicted terminal survival are different quantities; nevertheless, the relative stability of the model-predicted values across increasing severity is interpreted as evidence that the offline, model-based decision-support signal remains stable under severity shift.
To provide a secondary, physiology-focused assessment, the study introduces an episode-level physiological stabilization score (EPSS). EPSS is described as a heuristic summary that captures whether selected physiological variables move in favorable directions during follow-up. Using model-generated rollouts under offline policies, the authors compared EPSS values to matched logged clinical trajectories. For several physiological components, the model-generated rollouts received higher EPSS values than the matched clinical trajectories, indicating that simulated trajectories under the offline policies tended to show more favorable short-term physiological changes according to the EPSS heuristic.
The authors interpret the combination of results—the decline in observed clinical survival with increased severity alongside relatively stable, high model-predicted terminal survival and higher EPSS in model rollouts—as supporting learned-dynamics OPE as a useful severity-OOD stress test for offline RL policies in ICU sepsis. However, they explicitly note that learned-dynamics OPE and the simulation-based signals reported here are not a substitute for causal or prospective validation. The manuscript states that prospective and causal validation remain necessary next steps before offline RL-derived decision support could be considered for clinical use.
The paper is presented as a medRxiv preprint and has not been certified by peer review; the authors caution that the work should not be used to guide clinical practice at this stage. The authors also declare no competing interests and confirm adherence to ethical guidelines; the dataset used (MIMIC-III) requires credentialed access through PhysioNet.
All analyses were performed on MIMIC-III v1.4, available through PhysioNet under credentialed access. The authors state they cannot redistribute the patient-level data and direct interested researchers to PhysioNet for access after completing the required credentialing, training, and Data Use Agreement. The manuscript includes declarations that appropriate ethical approvals and consents were handled in accordance with relevant guidelines.
Note: This summary and the expanded discussion are based solely on the content reported in the preprint. The article is a preprint and has not undergone peer review; the source explicitly recommends that findings should not yet inform clinical practice without further validation.