Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

General-purpose large language models increasingly encounter emotional and therapy-like conversation, yet are not developed or evaluated as clinical systems. Existing safety evaluations rely largely on brief exchanges, although harms often unfold over extended interactions. Whether models maintain safety-relevant performance as conversations accumulate context remains unknown.

Methods

In this preregistered study, 400 clinician-validated statements, with or without suicidal ideation, were inserted at 0–200 speaker turns in 5 psychotherapy and 3 synthetic transcripts. Forty-nine LLMs and 8 clinicians performed the same binary classification task. Mixed-effects models estimated the effects of conversational depth, model scale, and model version on F1. Twelve top models were tested to 1,500 turns across conversational trajectories, with or without instruction restatement.

Results

F1 declined with depth across model families (p<0.001). Larger, newer models performed better but still degraded. Clinicians showed no decline (mean F1 0.86 at both 0 and 200 turns), but eight of nine proprietary models exceeded their performance at 200 turns. Conversational content, not length alone, explained F1 changes; the largest decrease was under adversarial context (p<0.001). Restating instructions increased F1 on therapy to near baseline (median ΔF1 +0.12; p<0.001; 89% median recovery) versus MSJ (ΔF1 +0.08; p=0.04; 38% recovery).

Conclusions

LLM detection of suicidal ideation degraded with conversational depth and trajectory, whereas clinician performance remained stable despite the strongest models exceeding most clinicians in absolute performance. Mental health AI safety evaluations should test sustained performance across realistic and adversarial trajectories rather than relying on short-prompt benchmarks.

Plain-language summary

Large language models increasingly encounter psychiatric distress and crisis-related content in ordinary use, but existing mental health AI safety benchmarks often rely on isolated prompts, brief exchanges, or simulated conversations rather than real extended interactions. In a preregistered benchmark of 1.4M inferences from 49 LLMs using real psychotherapy transcripts, suicidal-ideation detection declined with conversational depth in every model family while clinician performance remained stable — despite the strongest models matching or exceeding most clinicians in absolute performance — and, among top models, degradation at extended depths depended on the content and direction of the preceding conversation rather than length alone.

Article activity feed