Early identification of cognitive impairment is often limited in settings without access to comprehensive cognitive or clinical testing. Behavioural signals embedded in natural speech — including both acoustic and linguistic properties — are increasingly studied as potential markers of cognitive decline. The present proof-of-concept study evaluated whether integrating brief speech-derived measures with standard multi-domain cognitive testing improves discrimination between cognitively unimpaired individuals and those with amnestic mild cognitive impairment (aMCI) or early-stage Alzheimer’s disease.
The study combined features derived from one-minute naturalistic speech samples with results from a multi-domain cognitive assessment. The cognitive battery covered domains of attention, language, memory and executive function. Acoustic and linguistic features were extracted from each one-minute speech recording; the source describes these as complementary modalities but does not list specific features in the abstract.
Participants were classified into two groups for the primary analyses: cognitively unimpaired individuals and individuals with either aMCI or early-stage Alzheimer’s disease. The abstract reports that models were trained to distinguish between these diagnostic groups. Further participant demographics or sample size details were not provided in the source abstract.
Speech samples were processed to obtain two categories of speech-derived markers: acoustic features (relating to voice and sound characteristics) and linguistic features (relating to language content and structure). The study tested these speech-derived measures both alone and in combination with standard cognitive test scores to determine whether multimodal integration improved classification performance. Specific acoustic or linguistic variables and their extraction methods are not detailed in the abstract.
Multiple machine learning classification models were applied to compare performance across three feature sets: cognitive-only, speech-only (acoustic and linguistic), and combined cognitive plus speech features. The comparative approach assessed whether integrating modalities yields superior discrimination of cognitive impairment relative to single-modality models. The abstract reports statistical comparisons between model classes but does not specify which algorithms were used.
Across the machine learning models tested, combining cognitive, acoustic and linguistic features produced significantly better classification performance than models using cognitive features alone or speech features alone. The combined models achieved very high discrimination with reported area under the curve (AUC) values in the range of 0.96–0.98. Both comparisons (combined vs cognitive-only and combined vs speech-only) were reported as statistically significant (p < .05). The abstract frames these findings as evidence that multimodal integration enhances identification of cognitive impairment.
The investigators confirmed that all relevant ethical guidelines were followed. Ethical approval was granted by the Human Research Ethics Committees of QIMR Berghofer Medical Research Institute and The University of Queensland. The authors state that all necessary participant consent procedures were completed and that data produced in the study are available from the authors upon reasonable request. The manuscript is a preprint and has not been peer reviewed; the authors note that the report should not be used to guide clinical practice.
This proof-of-concept study suggests that integrating speech-based acoustic and linguistic measures with multi-domain cognitive assessment may substantially improve the accuracy of classifying early cognitive impairment. The authors propose that such multimodal approaches could support the development of accessible, scalable screening tools suited to primary care or other settings with limited resources. The abstract does not provide detailed information on sample size, model types, feature lists, external validation, or potential confounders, so additional methodological details and peer review will be needed to assess generalizability and clinical applicability.
The work was supported by the National Health and Medical Research Council (APP1135769). The authors declared no competing interests. As noted, this report is a preprint and has not undergone peer review; readers are advised not to change clinical practice based on the findings alone.
In this study, integrating one-minute speech-derived acoustic and linguistic features with standard cognitive testing improved machine-learning classification of cognitively unimpaired individuals versus those with aMCI or early Alzheimer’s disease, achieving AUCs of 0.96–0.98 (p < .05). The results provide a foundation for further work to validate multimodal, speech-informed screening tools that could be more accessible and scalable in primary care and community settings.