Large language model linguistic perplexity in childhood onset psychosis: unique features and developmental trends
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Objective
Child and early adolescent onset psychosis (COP) is associated with subtle changes in language linked to thought disorder, a key contributor to functional impairment. Large language models (LLM) can detect deviations from expected language patterns by jointly analyzing sentence structure and word choice. This integrated information is captured by measures such as (pseudo)-perplexity, which quantify how difficult it is for an LLM to predict individual words given the surrounding linguistic context. The objective was to determine if perplexity measures change with age and whether they are altered in COP.
Method
This study tested for a difference in perplexity in COP cases (N = 23) mean age 12.78 years as compared to controls (N = 15) mean age 11.67 years. Extensive manually transcribed interviews were analyzed (controls: 3,829; cases: 6,579 mean words).
Results
Perplexity derived from Large Language Model Meta AI (LLaMA) had a significant negative correlation with age in cases but not in controls. Pseudo-perplexity derived from Bidirectional Encoder Representations from Transformers (BERT) did not have a significant correlation with age in either group. Group differences were evaluated using generalized linear models with (pseudo)-perplexity as the dependent variable, case status as the predictor and age and number of words as covariates. The model predicting perplexity was significant and case status significantly predicted perplexity. In contrast, the model predicting pseudo-perplexity was not significant.
Conclusion
These differences in perplexity are interpreted as reflecting an altered developmental trajectory in the real-time semantic and syntactic planning that directs the flow of language in individuals with COP.
Plain language summary
Childhood onset psychosis (COP) is difficult to detect with subtle changes in language that impair communication and functioning. We used large language models (LLMs) to detect developmental changes in language in COP and controls. LLMs found language in COP harder to predict with improved prediction with increasing age in COP but not in controls. LLMs can use easily collected language samples to detect COP and identify changes in language that could be targeted to improve communication.