This preprint evaluated whether outputs from four large language models (LLMs) can discriminate experimentally manipulated pain-related cognitive states. Participants were randomly assigned to a negative, positive, or neutral pain-coping writing prompt and produced free-text responses during a 10-minute task. According to the report, ANOVA results supported discriminant validity: all four LLM-derived pain catastrophizing scores distinguished the negative coping condition from the positive and neutral conditions. This finding indicates that model-derived scores tracked the experimentally induced direction of coping-focused writing across models.
Convergent validity—correlation between model-derived scores and independent measures of state catastrophizing or pain—was model dependent. The authors report that only the Gemini 2.5 Pro-derived scores correlated with a post-challenge state PCS (r = 0.22) and with pain unpleasantness (r = 0.23). Other evaluated models (Claude Opus 4; GPT Mini 4o; Llama 4 Maverick) did not show the same pattern of correlation with the state PCS or pain unpleasantness as described in the source. These results suggest that at least one model produced outputs that aligned modestly with an established momentary measure of catastrophizing and with subjective unpleasantness.
Divergent validity was mixed in the reported analyses. LLM-derived scores were reported to be unrelated to pain intensity, which supports divergence from that specific pain dimension. However, small positive correlations with baseline trait catastrophizing were observed for two models: Gemini (r = 0.21) and Claude (r = 0.28), per the source. All LLM-derived scores also correlated with negative affect (reported rs ranging from 0.29 to 0.41). The magnitude of these correlations was similar to those observed for the state PCS, a pattern the authors interpret as evidence of limited specificity—LLM-derived catastrophizing outputs appear to reflect negative affect broadly as well as catastrophizing-related content.
The study sample included 91 adults with chronic pain who were receiving long-term opioid therapy; 57.3% were female and the mean age was 60.5 years. Participants completed baseline measures including the trait Pain Catastrophizing Scale (PCS). After random assignment, they completed a 10-minute writing task cued to a negative, positive, or neutral pain-coping condition. State affect and pain were assessed before and after the writing tasks and again after a cold pressor task (water at 4°C for up to 2 minutes). A state version of the PCS was administered after the cold pressor challenge. The source frames this as an approach to evaluate more ecologically valid, state-sensitive measurement of catastrophizing, in contrast to traditional trait-focused assessments.
Free-text responses were analyzed using four LLMs listed in the source: Claude Opus 4, GPT Mini 4o, Llama 4 Maverick, and Gemini 2.5 Pro. Each model produced a derived pain catastrophizing score from participants’ written responses; these model-derived scores formed the basis for ANOVA comparisons and correlational analyses with state and trait measures, pain intensity and unpleasantness, and affect. The preprint reports model-dependent differences in convergent and divergent validity, with Gemini demonstrating the most consistent modest correlations with state measures.
The authors present these findings as preliminary and explicitly note that further study is needed. The work is a clinical preprint and has not been peer reviewed; the source cautions that results should not guide clinical practice. IRB approval was provided by the Institutional Review Board at Stanford University and the trial is registered (NCT04097743). The source reports no competing interests. Scripts and LLM scores are reported to be hosted on OSF and available upon request when the link is public. Funding disclosed in the source includes support from the National Institute on Drug Abuse (T32DA035165 and K23DA048972).
According to the preprint, these results provide preliminary evidence that certain LLMs may serve as implicit markers of state pain catastrophizing when applied to free-text responses, but specificity is limited by correlations with negative affect and by model-dependent variability. The authors indicate that additional validation work is necessary, including replication, exploration of specificity versus affective confounds, and further comparison across models and tasks. The source does not report peer-review outcomes or additional replication data at the time of posting (July 24, 2026).