Sepsis remains a major cause of critical illness and death and presents ongoing diagnostic challenges. Whole-blood host-response transcriptomic profiling has been pursued for diagnosis and stratification, but implementation has been limited by cohort heterogeneity and evaluation design. The authors emphasize that AUROC (rank discrimination) and fixed-threshold decision behavior are distinct: a model may retain high AUROC while failing at a fixed decision threshold after cohort or platform transfer. This study therefore interrogates numerical transportability of decision thresholds rather than proposing a new biomarker signature.
The benchmark used four public whole-blood GEO cohorts assigned predefined roles: GSE65682 as discovery, GSE95233 as external microarray validation, GSE154918 as RNA-seq transfer validation, and GSE28750 as a non-infectious inflammation stress-test cohort. Inclusion required processed expression matrices and metadata enabling conservative phenotype harmonization. Analyses used only processed public matrices to preserve reproducibility and to avoid additional technical heterogeneity from raw-read reprocessing. Study size was determined by the availability of eligible public cohorts rather than a prospective sample-size calculation.
Four logistic-regression-based workflows were compared with respect to preprocessing and transfer:
The contrast is between strategies that require only training-set statistics for transfer (single-sample transfer) and approaches that use external-cohort distributional information (adaptation).
Internal performance was measured using five-fold cross-validation with fold-contained imputation and scaling to ensure that preprocessing within each fold did not leak test-set information. External uncertainty was estimated using 2,000 stratified bootstrap replicates for model-transfer evaluations to quantify variability in external metrics.
External validation involved applying the discovery-trained workflows to the predefined external cohorts. The primary emphasis was fixed external evaluation at a 0.5 decision threshold rather than ranking metrics alone. The authors reported balanced accuracy at that fixed threshold as a decision-focused metric and used bootstrap replicates to estimate uncertainty and compare paired workflows.
Across workflows, internal five-fold cross-validation showed very high discrimination (AUROC). However, external validation revealed a key distinction: training-derived scaling strategies (both standard and robust) suffered a collapse in fixed-threshold performance after transfer. In the RNA-seq transfer cohort (GSE154918), these training-derived approaches produced balanced accuracy of 0.50 at the 0.5 threshold despite very high AUROC—equivalent to random classification when using that fixed threshold.
This outcome illustrates that preserved rank ordering (high AUROC) does not guarantee preservation of the score scale required for direct application of a fixed threshold. From a translational standpoint, such hidden decision failure undermines the utility of a classifier when used as a single-sample, fixed-threshold diagnostic across cohorts or platforms.
The strict-inductive sample-wise rank normalization strategy preserved fixed-threshold performance across external cohorts in this benchmark. Reported balanced accuracy for this approach across transfer cohorts was high (reported range 0.95–1.00), indicating that rank-based single-sample preprocessing maintained score-scale stability adequate for applying a fixed 0.5 decision threshold after transfer. Because sample-rank normalization operates on individual samples without requiring external-cohort statistics, it is interpretable as a strict inductive single-sample transfer approach.
Robust scaling using unsupervised external-cohort reference statistics also preserved fixed-threshold performance (balanced accuracy 0.95–1.00 in the benchmark). The authors emphasize that this approach should be described as unsupervised cohort adaptation rather than single-sample transfer because it uses unlabeled external-cohort distribution information to estimate scaling parameters. Its practical reliability therefore depends on the availability of representative external reference samples at deployment time.
When the benchmark replaced healthy controls with a harder negative comparator (non-infectious inflammation, GSE28750), results shifted. Robust external-cohort adaptation achieved the highest observed balanced accuracy in this stress test (0.80, 95% CI 0.61–0.95). However, in paired bootstrap analysis the difference between robust adaptation and sample-rank normalization was uncertain, indicating that the advantage may not be robustly reproducible across resamples in this smaller or harder task.
The authors tested whether post hoc calibration or adjustments in model regularization could rescue failing training-derived scaling strategies. Those interventions did not restore usable fixed-threshold performance after transfer. This supports the interpretation that preprocessing and score-scale preservation before or during transfer are more critical for maintaining fixed-threshold decision behavior than post hoc calibration alone.
This benchmark demonstrates that high internal AUROC can mask decision failure at the level of fixed thresholds after cohort or platform transfer. Preprocessing choices that preserve the numeric score scale across cohorts — either by per-sample rank normalization or by unsupervised external-cohort scaling — substantially improved fixed-threshold transportability in this setting. Strategies that rely solely on training-derived scaling parameters were vulnerable to score-scale shifts that rendered a previously validated threshold ineffective.
The distinction between single-sample transfer and cohort adaptation is important for deployment: sample-rank normalization supports strict inductive single-sample use, while robust external-cohort scaling is best described as unsupervised adaptation and requires external reference samples for parameter estimation.
High AUROC does not guarantee that a classifier will retain usable fixed-threshold performance after cohort or platform transfer. In this multi-cohort whole-blood benchmark, strict-inductive sample-rank normalization was the most stable fixed external strategy, while robust external-cohort adaptation performed similarly but should be treated as an adaptation approach requiring unlabeled external reference samples. Post hoc calibration and regularization adjustments did not compensate for preprocessing-driven score-scale failure.
All source expression data and sample annotations are publicly available from GEO under accession numbers GSE65682, GSE95233, GSE154918, and GSE28750. Processed task matrices, benchmark outputs, figure source data, supporting tables, and analysis scripts are archived at Zenodo (DOI reported in the article) and mirrored in the public GitHub repository cited by the authors. The study used processed public matrices only and reports no funding or competing interests as declared by the authors.