Regression-based normative approaches are standard in neuropsychology but often depend on score transformations and distributional assumptions. This study directly compared traditional linear regression (LR)–based norms with GAMLSS (Generalized Additive Models for Location, Scale and Shape)–derived norms for a brief cognitive battery applied in the Norwegian Dementia Disease Initiation (DDI) cohort to determine whether improved modeling of score distributions materially changes clinical classifications and outcomes.
The analyses used the same normative samples that produced the original LR norms. Cognitive instruments examined were the CERAD word list test (including delayed recall), Trail Making Test (TMT) A and B, FAS phonemic fluency, and the Visual Object and Space Perception Battery (VOSP) Silhouettes. Two analytic cohorts were used: a normative subsample (n = 131) for expected low-score frequency and empirical base-rate assessment, and a clinical cohort from the DDI study (n = 643) to evaluate downstream clinical implications.
GAMLSS norms were developed using the same normative data inputs as the original LR norms. The comparison focused on (1) frequencies of expected low scores in the normative subsample, (2) concordance in classification between LR and GAMLSS in the clinical cohort, and (3) clinical associations including Mild Cognitive Impairment (MCI) classification, two-year diagnostic stability and change, and cerebrospinal fluid (CSF) biomarker profiles. Statistical agreement was quantified, and discordant classification rates were reported.
Compared to LR norms, GAMLSS produced lower frequencies of low scores overall. This reduction was primarily attributable to differences in the CERAD delayed recall distribution modeling. Despite these differences in low-score frequencies, overall classification concordance between approaches was high (kappa = 0.91). Only 4.2% of classifications were discordant between LR and GAMLSS, indicating substantial overlap in the practical identification of low performance across methods.
When applied to the DDI clinical cohort (n = 643), downstream outcomes—including MCI classification and two-year diagnostic stability and change—were broadly similar whether LR or GAMLSS norms were used. The modest rate of discordant classifications did not translate into large differences in longitudinal diagnostic trajectories at the two-year follow-up in this sample.
CSF biomarker profiles were evaluated in relation to cognitive classifications generated by each normative method. The study reports that biomarker results did not clearly favor either LR or GAMLSS norms; no decisive superiority of one normative approach emerged from the biomarker comparisons in this dataset.
The authors highlight that GAMLSS provides a more faithful representation of score distributions, especially for bounded and non-normal outcomes such as delayed recall measures. This improved fidelity led to fewer identified low scores in normative comparisons. However, because concordance with LR-based classification was high and clinical endpoints (diagnostic stability, MCI classification, and CSF associations) were similar, the clinical impact of switching from well-calibrated LR norms to GAMLSS was modest in this particular setting. The findings suggest that while GAMLSS offers statistical advantages in modeling, established LR norms that are properly calibrated may remain robust for many clinical classification tasks.
The DDI, TronderBrain, and Gothenburg MCI studies received approval from their respective regional or local research ethics committees. All participants provided written informed consent, and study procedures adhered to the Declaration of Helsinki and relevant national regulations. Datasets used in the analyses were de-identified prior to analysis; however, due to ethical restrictions the data generated for the study are not freely available for sharing.
Funding acknowledged in the report included the EU Joint Programme – Neurodegenerative Disease Research, The Research Council of Norway, and Helse Nord.
GAMLSS can better model complex, bounded, or non-normal neuropsychological score distributions and may reduce the frequency of low-score classifications for certain measures, notably CERAD delayed recall. Nevertheless, in this study the use of GAMLSS versus traditional linear regression norms produced only modest differences in clinical classification, two-year diagnostic outcomes, and CSF biomarker associations. The authors conclude that although GAMLSS offers statistical advantages, well-calibrated LR norms may remain appropriate and robust for clinical use in similar settings.