This study evaluated whether measures of linguistic unpredictability derived from Large Language Models change with age and whether they differ in children and early adolescents with childhood-onset psychosis (COP) compared with controls. The primary constructs were perplexity, reflecting how difficult it is for an LLM to predict a word in context, and pseudo-perplexity from a bidirectional transformer. The investigators framed these metrics as integrated indices of sentence-level syntactic and semantic planning that may reveal subtle language deviations linked to thought disorder in COP.
Participants comprised 23 COP cases (mean age 12.78 years) and 15 controls (mean age 11.67 years). All interviews were manually transcribed. The average amount of language analyzed per participant differed between groups: the mean words reported were 3,829 for controls and 6,579 for cases. The source reports these summary counts but does not provide participant-level distributions, sex breakdowns, clinical severity scales, medication status, or diagnostic subtypes in the abstract.
Two model families were used to generate prediction-based measures from transcript text. Perplexity was calculated using LLaMA (Large Language Model Meta AI), which yields a measure of how well the model predicts sequential words. Pseudo-perplexity was derived from BERT (Bidirectional Encoder Representations from Transformers), a bidirectional transformer that produces a related but distinct estimate of token predictability. The authors treat both metrics as reflecting joint information about word choice and sentence structure; however, they note differences between the two approaches in their sensitivity to developmental and diagnostic factors.
The analysis plan included correlation testing between age and the (pseudo)-perplexity measures separately within cases and controls. Group differences were evaluated using generalized linear models (GLMs) with (pseudo)-perplexity as the dependent variable, case status as the predictor, and age and number of words as covariates. The abstract reports which models were significant and whether case status predicted outcomes, but it does not include exact p-values, confidence intervals, or effect size estimates for these tests in the source abstract.
Key findings reported in the source are:
LLaMA-derived perplexity exhibited a significant negative correlation with age in COP cases but not in control participants. This indicates that within the COP group, older children had lower perplexity values derived from LLaMA.
BERT pseudo-perplexity did not show a significant correlation with age in either cases or controls.
In GLMs controlling for age and number of words, the model predicting LLaMA perplexity was significant and case status significantly predicted perplexity. In contrast, the GLM predicting BERT pseudo-perplexity was not significant.
The source does not report the exact statistical values (for example, correlation coefficients, regression coefficients, p-values, or model fit statistics) for these comparisons in the abstract; obtaining those details requires the full manuscript or contacting the authors.
The authors interpret the observed differences in LLaMA perplexity as evidence of an altered developmental trajectory in the real-time semantic and syntactic planning processes that direct language flow in individuals with COP. Specifically, the negative age–perplexity correlation in cases suggests a change with maturation that diverges from controls. The lack of age effects or group differences for BERT pseudo-perplexity highlights that sensitivity to developmental and diagnostic effects may vary across model architectures and metric definitions.
These interpretations are presented in the source as hypothesis-generating and linked to the concept that thought disorder and subtle language abnormalities contribute to functional impairment in early-onset psychosis.
The Rutgers University Institutional Review Board provided ethical approval, and the authors state that informed consent was obtained. Competing interests were declared as none. Summary-level data for the measures calculated in the paper will be made available on request; the original transcripts are not publicly shared.
Notable limitations based on the source abstract include the small sample sizes (23 cases, 15 controls), absence of reported inferential statistics (exact p-values and effect sizes were not given in the abstract), and limited detail on potential confounders such as medication, illness duration, symptom severity, socioeconomic or language background, and transcription or preprocessing steps. The source does not report whether analyses adjusted for additional clinical or demographic variables beyond age and number of words. These gaps indicate that readers should consult the full manuscript or contact the corresponding author for complete methodological and statistical detail.
Overall, the reported findings suggest that LLaMA-derived perplexity may capture developmental changes and diagnostic differences in language production in childhood-onset psychosis, whereas BERT pseudo-perplexity did not detect such patterns in this sample. Further work with larger samples and transparent reporting of effect sizes and analytic choices will be needed to confirm and extend these observations.