Detecting Self-Repairs from Spontaneous Speech with Prompt Ablation Across LLMs and Fine-Tuned Encoder
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Self-repairs – in-utterance revisions in which a speaker abandons and reformulates their speech – are a promising interpretable marker for speech-based cognitive screening. Detecting them automatically is difficult because a self-repair is defined by its relationship to surrounding speech rather than by fixed lexical cues. On the DementiaBank ADReSS corpus, we compared the capability of generative LLMs under a five-condition prompt ablation against a fine-tuned DistilBERT token classifier at detecting self-repairs. GPT-5 performed best (test F1 = 0.73) and was largely insensitive to prompt design, whereas the LlaMA (open-weight alternative) was both weaker and far more prompt-sensitive (test F1 = 0.47). DistilBERT, nearly 100 times smaller, matched the open-weight LLM at a fraction of the computational cost. These results suggest that a locally deployable encoder, given sufficient in-domain annotation, is a more plausible route to clinical self-repair detection than scaling model size or prompt complexity.