EEG-based machine-learning classifiers have shown promise for diagnosing neurodegenerative disorders, but their clinical translation depends on robustness to sample imbalance, center heterogeneity, and validation leakage. This multicenter study developed and applied a framework to evaluate diagnostic performance, calibration, and cross-center generalizability of EEG multifeatured classifiers across cognitively normal (CN), mild cognitive impairment (MCI), Alzheimer’s disease (AD), and frontotemporal dementia (FTD) samples, while explicitly addressing imbalance, statistical uncertainty, and validation rigor across six centers.
The analysis pooled data from six centers spanning several countries and institutions. Diagnostic groups included CN, MCI, AD, and FTD. Ethical approvals and informed consent were obtained from the institutional review boards or ethics committees at all participating sites listed in the source. The authors note that all participants provided written informed consent prior to participation.
The authors implemented supervised classifiers trained on multifeature EEG markers. To evaluate model performance and generalizability, they combined multiple validation schemes: aggregated- and subject-level repeated cross-validation and a leave-one-center-out (LOCO) approach. Validation rigor was increased by nesting calibration steps within strictly nested folds to prevent information leakage. The framework explicitly considered sample imbalance and statistical uncertainty when estimating performance metrics.
Calibration of predicted class probabilities was performed using Platt scaling, applied within nested cross-validation folds to avoid leakage. The framework evaluated both discrimination performance (e.g., AUC reported in the abstract) and the calibration/stability of predicted probabilities across validation regimes. The study assessed how well predicted probabilities correlated with clinical measures of cognitive impairment severity for each diagnostic contrast.
Across cross-validation and LOCO analyses, the CN versus AD contrast showed robust discrimination and consistent calibration. The abstract reports stable AUC and calibration across validation regimes and centers for this contrast. Predicted probabilities for AD were stable across validation schemes and correlated robustly with cognitive impairment severity, indicating that the classifiers produced both discriminative and calibrated outputs for AD versus cognitively normal participants.
Performance for CN versus MCI was described as moderate and heterogeneous across centers. In some cohorts, classification for this contrast reached only chance-level results, and overall cross-center generalizability was limited. For the MCI versus AD contrast, moderate discrimination was observed but only in the single available center where this comparison could be evaluated; broader generalizability for this contrast could not be established given limited cross-center data.
Contrasts involving frontotemporal dementia (FTD) showed modest or limited classification performance. The authors attribute these limited results primarily to sparse FTD samples across centers, which constrained both discrimination and calibration analyses for FTD-related contrasts. Because of sample scarcity, estimates of performance and generalizability for FTD remain uncertain.
Feature-importance analyses revealed disease-specific EEG signatures. For AD, characteristic markers included degradation of the alpha band and increases in slow-wave activity. In prodromal stages (MCI) and in differential contrasts involving FTD, patterns were weaker and more heterogeneous across centers. These findings suggest that while some EEG features consistently track AD-related neurophysiological changes, markers for MCI and FTD are less stable across heterogeneous samples.
The derived markers and computed features used for statistical analyses and model training are publicly available in a Zenodo repository, enabling reproducibility of the reported analyses. The source clarifies that raw EEG data are not publicly available due to ethical and legal restrictions and because they are owned by third-party institutions; researchers interested in raw data must contact the original data providers and comply with their access procedures.
The study concludes that EEG classifiers can provide robust discrimination and calibrated probability estimates for CN versus AD across multiple centers when validated with strict, leakage-avoiding methods. In contrast, classification for MCI and FTD was heterogeneous and showed limited cross-center generalizability, with some cohorts yielding chance-level performance. The authors emphasize the need for balanced sampling, standardized EEG and clinical protocols across centers, strict validation procedures including nested calibration, and explicit uncertainty quantification to support reliable clinical deployment of EEG-based dementia classifiers.
The preprint status indicates the work has not been peer reviewed. The dataset included sparse samples for some diagnostic groups (notably FTD), and some contrasts (e.g., MCI vs AD) were evaluated in only a single center. The raw EEG data are not publicly available due to ethical and ownership constraints. These limitations restrict the strength of claims about generalizability beyond the studied cohorts.
Researchers developing EEG diagnostic classifiers should prioritize balanced sampling across diagnostic groups, harmonized EEG acquisition and clinical classification protocols, and rigorous, leakage-free validation including nested calibration (for example, Platt scaling) and LOCO testing to assess cross-center performance. Reported discriminative performance for CN versus AD appears promising under these controls, but translation for MCI and FTD will require larger, more balanced multisite datasets and standardized procedures.