EEG foundation models (EEG-FMs) are commonly evaluated on their ability to discriminate disease states. The authors point out that for clinical biomarker use, measurement reliability—the stability of repeated measurements in the same individual—is a distinct requirement treated as a prerequisite by regulatory biomarker frameworks. The central question of this work is whether frozen EEG-FM representations provide reliable, repeatable measurements, whether reliability can be predicted from conventional model descriptors (such as pretraining paradigm or domain exposure), and what information in the representations supports reliability versus disease discrimination.
The analysis compared nine frozen representations: six EEG-focused foundation models spanning masked, contrastive, and predictive pretraining paradigms, handcrafted spectral features, and two general-purpose time-series models that had no EEG exposure during pretraining. All representations were evaluated under a single preprocessing pipeline across multiple datasets: two healthy retest cohorts (for test–retest reliability) and three neurodegenerative disease cohorts (for cross-sectional discrimination). The healthy retest cohorts provided retest intervals of roughly one month and roughly two years.
Reliability was quantified using the intraclass correlation coefficient (ICC) in healthy adults only. Across the nine representations, mean ICCs varied substantially, with values ranging from 0.08 to 0.76. The coefficient of variation (CV) for ICC across representations was reported as 53.0%, indicating large variability in stability across different frozen representations.
Disease-discrimination performance was measured cross-sectionally using area under the receiver operating characteristic curve (AUC) for three neurodegenerative cohorts. AUC values across the same nine representations showed much less dispersion than reliability measures; the AUC coefficient of variation was 5.4%. The authors report this as a roughly tenfold difference in relative dispersion between reliability and discrimination measures. They note that they did not conduct formal AUC equivalence testing and therefore describe discrimination as varying substantially less than reliability rather than claiming equivalence across models.
The observed variation in reliability and discrimination was not consistently explained by a model’s pretraining paradigm (masked, contrastive, predictive) or by whether the model was pretrained on EEG data. Notably, one of the general-purpose time-series models with no EEG exposure was among the most reliable representations tested. The lack of a consistent explanatory relationship means that reliability cannot be inferred reliably from typical model descriptors and therefore must be measured directly when considering models for longitudinal or biomarker applications.
The authors examined which frequency bands contributed to the two properties. Alpha-band information contributed disproportionately to measurement reliability, whereas theta-band information ranked first for discriminating Alzheimer disease and frontotemporal dementia. These band-specific findings are directionally consistent with established EEG evidence for these conditions. The reported statistical significance applied to the reliability contrast; the disease-related contrasts were described in directional terms aligned with prior literature.
Two methodologically distinct analyses—subspace geometry and band-ablation—were applied to probe how representations encode reliability and disease sensitivity. Both analyses indicate that reliability and discrimination are partially, not fully, dissociable properties of EEG-FM representations. The authors further report that network architecture influences whether dissociation between reliability and discrimination is preserved across model depth, traded off between properties, or leads to joint degradation of both.
Because reliability showed no detectable association with discrimination performance, pretraining paradigm, or domain exposure in the models studied, the authors recommend that reliability become a standard evaluation axis for EEG-FM representations intended for longitudinal monitoring or biomarker development. They emphasize that reliability must be measured directly rather than inferred.
The authors also released a reproducible pipeline for their evaluations to facilitate replication and adoption of reliability assessment in future EEG-FM work.
The source lists several OpenNeuro datasets used in the work. The authors declared no competing interests. Specific dataset identifiers referenced in the source include https://openneuro.org/datasets/ds004148, ds007176, ds004504, ds002778, and ds003490. The source also indicates that the authors released their reproducible analysis pipeline, but details of repository locations or file-level descriptions were not reported in the provided text.