Assessing Pain Catastrophizing Through Free-Text Responses: A Validation of Large Language Models

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Validated measures of pain catastrophizing primarily assess catastrophizing as a stable trait. However, emerging evidence suggests catastrophizing fluctuates with context, highlighting a need for ecologically valid methods to capture it. This study evaluated large language models (LLMs) as implicit markers of catastrophizing from free-text responses from ninety-one adults with chronic pain receiving long-term opioid therapy (57.3% Female; mean age = 60.5 years). Patients completed baseline measures, including the trait pain catastrophizing scale (PCS), followed by a 10-minute writing task after random assignment to a negative, positive, or neutral pain-coping condition. State affect and pain were assessed before and after writing tasks and again after a cold pressor task (4ºC; ≤ 2 minutes). A state PCS followed the cold pressor task. Free-text responses were analyzed using four LLMs (Claude Opus 4; GPT Mini 4o; Llama 4 Maverick; and Gemini 2.5 Pro). ANOVA-based results supported discriminant validity, as all four LLM-derived pain catastrophizing scores differentiated negative from positive and neutral pain-coping conditions. Convergent validity was model-dependent; only Gemini-derived scores correlated with state catastrophizing (r = .22) and pain unpleasantness (r = .23). Divergent validity was mixed. LLM-derived scores were unrelated to pain intensity, but Gemini and Claude-derived scores showed small correlations with trait PCS (r’s = .21; 28, respectively). All LLM-derived scores also correlated with negative affect (r’s = .29–.41), comparable in magnitude to state PCS, suggesting limited specificity. These findings provide preliminary evidence that certain LLMs may serve as implicit markers of state pain catastrophizing, but further study is needed.

Summary

Leading Large Language Models (LLMs) demonstrate discriminant validity but mixed convergent and divergent validity scoring pain catastrophizing from persons living with chronic pain free-text responses.

Article activity feed