The authors report a failure mode they term rejection collapse, where standard uncertainty-informed rejection procedures unexpectedly produce a severe drop in global predictive performance. This collapse reflects localized vulnerabilities that are not apparent from common aggregate machine learning metrics. To systematically diagnose the phenomenon, the study uses a clinical prediction task as a proof of concept and applies a pipeline combining ensemble modeling, uncertainty decomposition, and unsupervised subgroup discovery.
As a demonstrator application, the team used prediction of Levodopa-Induced Dyskinesia in Parkinson's disease. The preprint frames this clinical task as a suitable setting to reveal atypical uncertainty-driven failures because it presents real-world heterogeneity and potential data ambiguity across patient subgroups. The manuscript emphasises that the experimental work is a proof of concept for diagnostic methodology rather than a clinically validated predictive tool.
To characterise prediction and confidence, the researchers trained a heterogeneous machine learning ensemble. They decomposed predictive uncertainty into Aleatoric and Epistemic components to separate irreducible data noise from uncertainty due to model knowledge. This decomposition was central to interrogating whether confident errors were associated with inherent ambiguity in the data or with lack of model knowledge.
After obtaining ensemble predictions and uncertainty estimates, the authors applied unsupervised subgroup discovery to the prediction set. This step aimed to locate subpopulations with distinct error and uncertainty patterns that a global evaluation would mask. Stratified error analysis across discovered subgroups revealed divergent predictive regimes that explained the global performance collapse.
The analysis identified two contrasting subgroups. In one subgroup, the models extracted a clear predictive signal and performed as expected. In the other subgroup, baseline features lacked discriminative capacity; this cohort produced a high rate of confident misclassifications. Crucially, these confident errors occurred below the system's rejection thresholds, meaning the rejection mechanism did not filter them out. Operating below rejection cutoffs, the problematic subgroup suppressed aggregate predictive metrics and produced the observed global rejection-curve collapse.
The authors emphasise that the root cause was subgroup-specific data ambiguity rather than a failure of the uncertainty-estimation algorithm itself. In other words, the ensemble and decomposition exposed where data did not support reliable discrimination, and that absence of signal led to confidently wrong predictions that escaped rejection.
From their findings, the authors argue that localized, subgroup-aware evaluation is a methodological requirement before deploying uncertainty-aware models in real-world clinical settings. Global rejection curves and aggregate metrics can conceal cohort-level vulnerabilities; without subgroup stratification and uncertainty decomposition, teams may deploy models that appear robust overall but fail catastrophically on specific cohorts.
The study implies several practical takeaways:
These recommendations flow from the proof-of-concept demonstration rather than prescriptive clinical guidance; the authors note the limits of a preprint-stage methodological investigation.
The preprint states that all individual-level data were fully de-identified prior to use and that relevant ethical approvals and participant consents were obtained. The authors declare no competing interests. All datasets and code supporting the findings are available from the authors on reasonable request, according to the Data Availability statement.
The manuscript is a medRxiv preprint and has not undergone peer review. The authors explicitly state that the research should not be used to guide clinical practice until it has been validated through standard peer review and further work.