---
title: "High AUROC can hide threshold failure in sepsis transcriptomic classifiers: preprocessing stabilit"
id: "plos-one-13-high-auroc-can-mask-decision-failure-in-sepsis-transcriptomic-classifiers"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-13-high-auroc-can-mask-decision-failure-in-sepsis-transcriptomic-classifiers"
content_type: "clinical_feed_article"
specialty: "Infectious Disease"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357585"
published_at: "2026-09-03T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# High AUROC can hide threshold failure in sepsis transcriptomic classifiers: preprocessing stabilit
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-13-high-auroc-can-mask-decision-failure-in-sepsis-transcriptomic-classifiers
- **Specialty:** [Infectious Disease](https://medichelpline.com/clinical-feed/infectious-disease.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357585)
- **Published At:** 2026-09-03T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- The study benchmarks numerical transportability of whole-blood **sepsis** transcriptomic classifiers across four public GEO cohorts: GSE65682 (discovery), GSE95233 (microarray validation), GSE154918 (RNA-seq transfer), and GSE28750 (non-infectious inflammation stress test). - The authors compared four preprocessing/model-transfer workflows: training-derived standard scaling, training-derived robust scaling, strict-inductive sample-wise **rank normalization** with training-derived scaling, and robust scaling using unsupervised external-cohort reference statistics (termed external-cohort adaptation). - Internal five-fold cross-validation showed very high discrimination (AUROC) for all workflows, but AUROC did not reliably predict fixed-threshold performance after cohort or platform transfer. - In external validation (GSE154918 RNA-seq transfer), training-derived standard and robust scaling experienced threshold collapse: balanced accuracy at the 0.5 decision threshold fell to 0.50 despite very high AUROC, equivalent to random classification at that fixed threshold. - The strict-inductive sample-rank strategy preserved fixed-threshold performance across external cohorts (balanced accuracy 0.95–1.00 in the benchmark), indicating strong score-scale stability for single-sample external transfer. - Robust external-cohort adaptation also preserved fixed-threshold performance (balanced accuracy 0.95–1.00) but requires access to unlabeled external-cohort distribution statistics and is therefore framed as unsupervised cohort adaptation rather than pure single-sample transfer. - Post hoc calibration and variations in model regularization did not rescue the failing training-derived-scaling strategies in external transfer. - In the sepsis-versus-non-infectious-inflammation stress test, robust external-cohort adaptation achieved the highest observed balanced accuracy (0.80, 95% CI 0.61–0.95), but its advantage over sample-rank normalization was uncertain in paired bootstrap comparisons. - The authors conclude that high **AUROC** can mask fixed-threshold decision failure after transfer; preprocessing stability (sample-rank or external-cohort scaling) outweighs post hoc calibration for preserving usable fixed thresholds across cohorts and platforms. - All analyses used processed public expression matrices only; data and scripts are archived on Zenodo and GitHub and cohorts are available under the reported GEO accession numbers.
## Clinical Analysis & Structured Key Points
High AUROC can mask decision failure in sepsis transcriptomic classifiers: Preprocessing stability outweighs post hoc calibration across cohorts | PLOS One Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click here . Article Authors Metrics Comments Media Coverage Reader Comments Figures Figures Abstract Background High AUROC is often taken as evidence that a transcriptomic classifier is promising, but rank discrimination can conceal fixed-threshold failure after cohort or platform transfer. Methods We benchmarked four GEO whole-blood cohorts: GSE65682 for discovery, GSE95233 for external microarray validation, GSE154918 for cross-platform RNA-seq validation, and GSE28750 for non-infectious inflammation stress testing. We compared logistic-regression workflows using training-derived standard scaling, training-derived robust scaling, sample-wise rank normalization with training-derived scaling, and robust scaling using unsupervised external-cohort reference statistics. Internal performance used five-fold cross-validation with fold-contained imputation and scaling. External uncertainty used 2,000 stratified bootstrap replicates. Results Internal discrimination was very high for all strategies, but external validation revealed threshold collapse for training-derived standard and robust scaling. In GSE154918, both had balanced accuracy 0.50 at the 0.5 threshold despite very high AUROC, equivalent to random classification at that fixed threshold. The strict-inductive sample-rank strategy preserved fixed-threshold performance across external cohorts (balanced accuracy 0.95–1.00). Robust external-cohort adaptation also performed well (0.95–1.00) but uses unlabeled external-cohort distribution statistics and is therefore reported as adaptation rather than fixed single-sample transfer. Calibration and regularization sensitivity did not rescue the failing training-derived scaling strategies. In the sepsis-versus-non-infectious-inflammation stress test, robust external-cohort adaptation had the highest observed balanced accuracy (0.80, 95% CI 0.61–0.95), but its difference from sample-rank normalization was uncertain in paired bootstrap analysis. Conclusions High AUROC can mask fixed-threshold failure in sepsis transcriptomic classifiers. In this benchmark, strict-inductive sample-rank normalization was the most stable fixed external strategy, while robust external-cohort scaling was best interpreted as unsupervised cohort adaptation whose reliability depends on external reference-sample availability. Citation: Zheng H, Chen W (2026) High AUROC can mask decision failure in sepsis transcriptomic classifiers: Preprocessing stability outweighs post hoc calibration across cohorts. PLoS One 21(9): e0357585. https://doi.org/10.1371/journal.pone.0357585 Editor: Vijayalakshmi Kakulapati, Sreenidhi Institute of Science and Technology, INDIA Received: March 17, 2026; Accepted: August 19, 2026; Published: September 3, 2026 Copyright: © 2026 Zheng, Chen. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: All source expression data and sample annotations analyzed in this study are publicly available from the Gene Expression Omnibus (GEO) under accession numbers GSE65682, GSE95233, GSE154918, and GSE28750. The processed task matrices, benchmark outputs, figure source data, supporting tables, and analysis scripts underlying the reported findings are archived at Zenodo ( https://doi.org/10.5281/zenodo.21850714 ) and are mirrored in the public GitHub repository ( https://github.com/Zhenghongwei11/High-AUROC-sepsis-decision-failure,releasetagv1.2.0-reproducibility ). There are no restrictions on access beyond those imposed by the original public repositories. Funding: The author(s) received no specific funding for this work. Competing interests: The authors have declared that no competing interests exist. Introduction Sepsis is defined as life-threatening organ dysfunction caused by a dysregulated host response to infection and remains a major global cause of death and critical illness [ 1 , 2 ]. Because early recognition remains difficult, whole-blood host-response profiling has attracted sustained interest as a route to diagnosis-oriented and stratification-oriented research [ 3 ]. At the same time, biomarker implementation in sepsis has remained difficult, in part because cohort heterogeneity, comparator choice, and evaluation design can all distort apparent performance [ 3 ]. One recurring problem is that transcriptomic models are often developed under analytically permissive conditions. Early molecular diagnostic work showed that host-response assays can separate sepsis-related states under defined development and validation settings [ 4 ]. More recent public-data machine-learning studies have also used compact signatures and internal or discovery-stage validation to nominate sepsis biomarkers or prognostic models [ 5 , 6 ]. Those designs may still be useful for hypothesis generation, but they do not answer the more practical question of whether a fixed workflow remains numerically stable after cohort shift, platform shift, or replacement of healthy controls with harder inflammatory comparators. This distinction matters because rank discrimination and threshold behavior are not the same property. A workflow may preserve AUROC while still failing at the level where a classification threshold would be applied. In that setting, strong internal discrimination and strong external ordering can coexist with an unstable external score scale. From a translational perspective, that hidden decision failure is more important than another increment in internal cross-validation performance because externally transferred models are interpreted at the level of fixed scores and thresholds, not at the level of retrospective ranking alone. Multicentre host-response studies with explicit validation demonstrate that blood RNA measurements can carry reproducible infection-related signal [ 7 , 8 ]. The unresolved gap is numerical transportability. Sepsis biomarker guidance emphasizes the difficulty of moving promising markers into robust implementation [ 3 ]. In parallel, public transcriptomic machine-learning studies often prioritize biomarker nomination or prognostic modeling rather than direct testing of whether fixed scores and thresholds remain usable after transfer [ 5 , 6 ]. Biomarker discovery and workflow evaluation are therefore often conflated. We therefore interrogated the numerical transportability of classification thresholds rather than proposing another sepsis signature. Using a multi-cohort whole-blood GEO benchmark, we tested whether AUROC can overstate classifier usefulness once evaluation moves from internal validation to fixed external transfer. We intentionally used conservative phenotype harmonization, gene-level feature alignment, simple baseline models, and fixed external evaluation so that transportability failures would remain interpretable rather than being obscured by modeling flexibility. The core hypothesis was that externally stable transcriptomic classification depends more on decision-score stability under transfer than on nominal internal discrimination, and that preprocessing rigor matters more than post hoc calibration or nominal model complexity for preserving that stability. We also expected that workflow ranking would change once the benchmark moved from biologically distant healthy controls to harder inflammatory negative comparators. This design follows the logic of transparent prediction-model reporting, where the credibility of a model study depends not only on discrimination but on how clearly the development and evaluation process has been made visible [ 9 , 10 ]. Materials and methods Study design This study evaluated how phenotype harmonization, preprocessing stability, and validation design influence the transportability of whole-blood transcriptomic classifiers for sepsis-related tasks using public GEO cohorts. The emphasis was on robustness of an existing workflow under external transfer rather than de novo signature discovery [ 9 , 10 ]. That emphasis is aligned with guidance that prediction studies should report discrimination and calibration together rather than rely on AUROC alone [ 11 , 12 ]. Decision-curve and risk-of-bias frameworks further motivate separating threshold behavior, applicability, and model-evaluation context from headline discrimination metrics [ 13 , 14 ]. A study-design overview is provided in Fig 1 . Download: PNG larger image TIFF original image Fig 1. Study design and analytical framework of the transportability benchmark. Panel A summarizes the four public whole-blood cohorts and their assigned analytical roles: GSE65682 for discovery, GSE95233 for external microarray validation, GSE154918 for RNA-seq transfer validation, and GSE28750 for inflammatory-comparator stress testing. Panel B outlines the benchmark structure, including the primary sepsis-or-shock-versus-healthy analysis (n = 275), the sepsis-versus-non-infectious-inflammation analysis (n = 147), and the nested RNA-seq transfer readout. Panel C summarizes the workflow comparison from preprocessing to fixed external validation and decision-focused metrics. Panel D illustrates the interpretive framework: high AUROC can coexist with fixed-threshold failure after transfer when score scale is not preserved. https://doi.org/10.1371/journal.pone.0357585.g001 All analyses were restricted to processed public expression matrices. This was done deliberately to keep the workflow reproducible on modest local compute and to avoid raw-read reprocessing as an additional untracked source of technical heterogeneity. Study size was determined by availability of eligible public cohorts rather than by a prospective sample-size calculation. We therefore treated this as a convenience benchmark built from all public cohorts that met the eligibility criteria defined for this retrospective benchmark. Public cohorts and analytical roles Four public whole-blood cohorts were assigned analytical roles before model benchmarking: GSE65682 as the primary discovery cohort, GSE95233 as the external microarray validation cohort, GSE154918 as the RNA-seq transfer validation cohort, and GSE28750 as the non-infectious inflammation stress-test cohort. Cohorts were included only if processed matrices were publicly available and their metadata supported conservative harmonization into phenotype label groups. A summary of cohort roles, retained labels, included and excluded samples, and analysis membership is provided in Table 1 . Download: PNG larger image TIFF original image Table 1. Cohort roles, phenotype-harmonization counts, and analysis membership for the public whole-blood benchmark cohorts. The table summarizes GEO accession, platform type, total series size, samples retained after conservative phenotype harmonization, excluded samples, labels available after harmonization, assigned analytical role, task membership, and gene-level feature counts. Retained counts describe samples kept in the unified metadata table after phenotype harmonization and expression-matrix availability checks; task-membership counts describe the smaller binary-contrast subsets used for each benchmark analysis. Feature counts refer to gene-level rows retained after platform-specific probe or gene-symbol harmonization. https://doi.org/10.1371/journal.pone.0357585.t001 After phenotype harmonization and first-pass inclusion filtering, GSE65682 contributed 359 analyzable samples to the unified metadata table, GSE95233 contributed 73, GSE154918 contributed 91, and GSE28750 contributed 41. The larger retained count for GSE65682 reflects all harmonized labels available for the unified metadata table; only the subset matching each binary contrast entered a given benchmark task. Thus, 93 GSE65682 samples entered the primary sepsis-or-shock-versus-healthy training analysis, whereas infection-but-not-sepsis and non-infectious-inflammation samples were retained for metadata traceability or for the secondary comparator analysis. Several source cohorts were assembled before Sepsis-3 and do not support retrospective relabeling to a single modern consensus definition from processed GEO records alone. We therefore harmonized labels against the original study descriptions and recoverable processed metadata, and we report GEO accessions, platform types, and benchmark roles rather than incomplete cohort-date fields. To distinguish genuinely unavailable variables from merely unmodeled ones, we also performed a cohort-level metadata completeness audit across age, sex, infection or clinical context, severity or outcome, timepoint or follow-up, and treatment or exposure fields (S3 Table in S1 File ). Phenotype harmonization Phenotype harmonization was intentionally conservative. Samples were mapped into a compact label dictionary consisting of healthy, infection_non_sepsis, sepsis, septic_shock, and noninfectious_inflammation. Samples lacking defensible metadata support for this mapping were excluded rather than forced into pooled labels. For GSE95233, only healthy controls and day-1 septic-shock samples were retained, while later follow-up septic-shock samples were excluded. For GSE154918, Hlty, Inf1_P, Seps_P, and Shock_P samples were retained, whereas follow-up samples were excluded. For GSE28750, all healthy, sepsis, and post-surgical samples were retained. For GSE65682, only healthy subjects and ICU samples with metadata signals consistent with abdo_s, ctrl_GI, cap, hap, or no-cap were retained for the present benchmark. This strategy prioritized label defensibility over maximal sample count because the benchmark objective was workflow stability under realistic but interpretable phenotype definitions. The main risk of this decision is reduced sample size; the benefit is a cleaner comparator structure for external failure-mode analysis. Expression matrix extraction and gene-level harmonization Per-cohort included-sample expression matrices were extracted from GEO series matrices or, where required, from GEO supplementary processed files. Array-based cohorts were then collapsed to the gene-symbol level using GEO platform annotation tables. Probe rows were retained only if the platform Gene Symbol field mapped conservatively to a single symbol. Probe rows with blank annotation or multiple mapped symbols were excluded. When multiple probe rows mapped to the same gene, they were collapsed using the cohort-level median rather than selecting the highest-variance or best-performing probe. RNA-seq data from GSE154918 were harmonized at the gene-symbol level from processed matrix annotation columns and were collapsed by the same median rule when duplicate gene symbols were present. This produced gene-level matrices for all four cohorts. The resulting harmonization statistics were as follows: GSE154918 retained 19,203 gene rows; GSE28750 and GSE95233 retained 21,815 gene rows each; and GSE65682 retained 11,222 gene rows. The smaller GSE65682 feature space reflects conservative single-symbol probe mapping rather than downstream model-stage data loss. Shared gene spaces were defined at the task level by intersecting retained gene symbols before model fitting. This operation used feature availability only and did not use outcomes, predicted scores, external-cohort labels, or external model performance. The gene intersection was therefore fixed for each task before internal cross-validation began rather than reselected inside each fold. Predefined benchmark analyses Two benchmark analyses were defined before modeling, with cross-platform transfer treated as a predefined property of the primary analysis rather than as a separate discovery branch. The primary analysis evaluated sepsis or septic shock versus healthy controls across four cohorts. This analysis contained 275 samples and 9,508 shared genes. Within it, GSE65682 contributed 93 training samples, GSE95233 contributed 73 external-validation samples, GSE154918 contributed 79 RNA-seq transfer-validation samples, and GSE28750 contributed 30 additional external stress-test samples. The GSE154918 subset was interpreted as the cross-platform robustness readout for the primary analysis. The secondary analysis evaluated sepsis versus noninfectious_inflammation. This analysis contained 147 samples and 9,984 shared genes. GSE65682 served as the discovery cohort with 126 samples, and GSE28750 served as the external stress-test cohort with 21 samples. The cohort roles, primary contrast, secondary contrast, and RNA-seq transfer readout were defined before model fitting. Threshold sweeps, paired bootstrap differences, calibration diagnostics, regularization sensitivity, coefficient-stability summaries, and adaptive-reference-size sensitivity were added during revision as diagnostic or sensitivity analyses to clarify the observed failure modes. Workflow factors under comparison The first benchmarking stage focused on preprocessing and external-reference strategy rather than on proliferation of model families. Four strategies were compared so that strictly inductive transfer and unsupervised external-cohort adaptation were separated explicitly. Standard scaling (train-only) used per-feature median imputation followed by feature-wise standard scaling. For internal cross-validation, the imputer and scaler were fitted within each training fold and applied to the held-out fold. For external evaluation, the imputer and scaler were fitted once on the full GSE65682 training cohort and applied unchanged to each external cohort. Robust scaling (train-only) used per-feature median imputation followed by median and interquartile-range scaling with all reference statistics estimated from the training fold or training cohort only. This strategy represents the strict-inductive counterpart to robust scaling. Sample-wise rank normalization first replaced each sample’s expression values with within-sample percentile ranks and then applied feature-wise standard scaling with training-derived parameters. Because the rank step is sample-wise, external transformation does not require information from other external samples. Adaptive robust scaling (external reference) used the same training-cohort robust scaling for model fitting, but external samples were transformed using medians and interquartile ranges estimated from the unlabeled external cohort. This strategy does not use external class labels, but it does use the external-cohort feature distribution. We therefore treat it as unsupervised external-cohort adaptation or transductive preprocessing, not as a strictly inductive fixed-model workflow for individual external samples scored one at a time. These strategies encode different assumptions about how an external cohort should be aligned to the score axis learned in the training cohort. Standard z-score preprocessing anchors each feature to the training-cohort mean and standard deviation. Training-only robust scaling changes the reference estimator but remains fixed after training. Sample-wise rank normalization reduces dependence on absolute measurement scale by replacing each sample with within-sample order statistics before feature scaling. Robust external-cohort adaptation uses unlabeled external-distribution information to estimate a new reference frame. The comparison was motivated by foundational work showing that batch effects can dominate high-throughput expression analyses [ 15 , 16 ]. Cross-study normalization and platform-reproducibility studies further show that expression measurements can require explicit alignment before model transfer [ 17 – 19 ]. The external-adaptation arm was also inform
## Related Clinical Research

- [Transcriptome differences linked to delayed mortality in invasive group A Streptococcal (iGAS) dis](https://medichelpline.com/clinical-feed/plos-one-13-transcriptome-profile-of-delayed-mortality-in-patients-with-invasive-group-a.md)
- [Iron-related Parameters and Sepsis-associated Acute Kidney Injury — Source Content Not Available](https://medichelpline.com/clinical-feed/frontiers-in-immunology-19-association-of-iron-related-parameters-with-the-development-of-sepsis.md)
- [Mitochondria-associated genes and sepsis: source article content not available](https://medichelpline.com/clinical-feed/frontiers-in-immunology-15-potential-mitochondria-associated-pathogenic-genes-in-sepsis-a-multi-omics.md)
- [Trauma immune response: hallmarks, determinants and precision immunomodulation in humans](https://medichelpline.com/clinical-feed/nature-immunology-5-hallmarks-of-the-multidimensional-immune-response-to-trauma-in-humans.md)
- [Wickerhamomyces anomalus Pseudo-Outbreak Linked to Contaminated Sterile Lubricating Gel in South A](https://medichelpline.com/clinical-feed/cdc-emerging-infectious-diseases-journal-2-wickerhamomyces-anomalus-pseudo-outbreak-associated-with-medical-lubricating.md)

## Navigation
- [← Back to Infectious Disease Feed](https://medichelpline.com/clinical-feed/infectious-disease.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.