Sepsis is a leading cause of hospital mortality but is difficult to identify early because presentations are nonspecific and code-based case definitions suffer from label noise. To address this, the authors developed STRIDE, a machine-learning framework designed for scalable and accurate sepsis detection across multiple hospitals. The central innovation was a pragmatic, scalable approach to improve label quality by using a large language model (LLM) applied to discharge summaries, with an independent physician-adjudicated cohort defined as the gold standard for validation.
The analysis included a multi-institutional dataset covering seven hospitals and a total of 356,610 inpatient encounters. The work received ethical oversight from the Hartford HealthCare Institutional Review Board. The authors note that access to the underlying data requires appropriate regulatory approvals, including HIPAA training or certification and authorization by relevant institutional review boards and the collaborating healthcare system.
To mitigate the limitations of code-based sepsis labels, the investigators refined a pragmatic operational definition by applying a large language model to discharge summaries. This LLM-derived refinement was used to improve label quality at scale. For an independent assessment of label validity, the team assembled a physician-adjudicated cohort that served as the gold-standard reference for model validation. The source reports that this combination of automated LLM-assisted refinement and manual physician adjudication underpinned the evaluation of STRIDE.
STRIDE was evaluated across three observation-window durations commonly considered in sepsis prediction work: 8 hours, 24 hours, and 48 hours. Models were trained and derived from the multi-hospital dataset, and performance was assessed both in derivation and against the physician-adjudicated validation cohort. The authors highlight that the approach requires minimal prior patient history, supporting deployment in settings where longitudinal data are limited.
Across the full sample of 356,610 encounters, the 8-hour STRIDE model achieved an area under the receiver operating characteristic curve (AUC) of 0.960 in the derivation set. When evaluated against the independent physician-adjudicated validation cohort, the 8-hour model achieved an AUC of 0.878.
The study compared STRIDE against established sepsis identification benchmarks: SOFA, SIRS, and Epic's sepsis detection approach. STRIDE outperformed SOFA and Epic on discrimination metrics and demonstrated favorable calibration as measured by the Brier score. The authors reported that STRIDE retained strong discrimination among encounters that were SIRS-positive but adjudicated as non-septic, indicating specificity in distinguishing true sepsis from non-septic systemic inflammatory responses.
In the physician-adjudicated validation cohort, the STRIDE 8-hour model reached 78.2% specificity at a sensitivity of 80%, a performance characteristic the authors emphasize as balancing timely detection with limiting unnecessary alerts.
The study explored model behavior across different observation windows and highlighted that the 8-hour model matched or outperformed the longer 24- and 48-hour models in both derivation and validation. This suggests that shorter-window detection can achieve high discrimination while reducing the need for extensive prior data. The reported specificity at a set sensitivity threshold suggests STRIDE can limit false alerts while maintaining clinically meaningful sensitivity, an important consideration for surveillance and clinical decision support.
The authors note that access to the study data is controlled and requires institutional and regulatory approvals. The report indicates ethical oversight and approvals were obtained and that patient consent procedures and reporting guidelines were followed. As this work is presented as a medRxiv preprint, readers should interpret the findings in that context. The source document provides no additional methodological detail on model architecture, feature sets, or the specific large language model used beyond stating that an LLM was applied to discharge summaries for label refinement.
STRIDE represents a machine-learning approach that combines language model–assisted label refinement with physician adjudication to improve sepsis detection from electronic health records across multiple hospitals. In this study population of 356,610 encounters, the 8-hour STRIDE model showed strong discrimination (AUC 0.960 derivation; AUC 0.878 physician-adjudicated validation), favorable calibration, and better discrimination than SOFA and Epic benchmarks. The framework aims to support accurate sepsis surveillance while limiting unnecessary alerts and minimizing dependence on extensive prior history.