Automatic Speech Recognition and Phonetics-Informed Sentence Design for Spastic Dysarthria Detection and Corticobulbar Lesion Localization

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Spastic dysarthria diagnosis through subjective neurologist auditory-perceptual assessment remains standard practice despite known inaccuracy. To address this gap, we developed an objective framework grounded in phonetic evidence that spastic dysarthria preferentially impairs initial consonant articulation, using automatic speech recognition (ASR) to quantify dysarthria and localize corticobulbar lesions. We created four reading sentences targeting groups of initial consonants: labial (facial), lingual-alveolar (tongue), and velopharyngeal (pharyngeal/soft-palate) sentence, along with a mixed-consonant sentence for comparative evaluation. Thirty-seven patients with neuroimaging-confirmed corticobulbar lesions and 37 controls read each sentence. ASR transcribed dysarthric speech into text, and we computed a ‘syllable-error score’ by counting incorrectly transcribed syllables. This yields a clinically meaningful feature that makes syllable-level phonetic errors explicit. Logistic regression models were trained for each sentence, and performance was summarized by the area under the receiver operating characteristic curve (AUC) across 10,000 resampled train-test splits. Consonant-specific sentences significantly outperformed the mixed sentence: the lingual-alveolar sentence performed best with (median AUC 0.88), followed by the labial (0.80), then the velopharyngeal sentence (0.72), while the mixed-consonant sentence was lowest (0.67). These results suggest that the interpretable ASR-derived syllable error feature, combined with a relevant machine learning classifier could inform clinical insight into consonant-specific vulnerability in spastic dysarthria, with lingual-alveolar consonants appearing particularly informative. Overall, this novel ASR-based framework, together with phonetics-informed feature design provides objective, accurate, and clinically meaningful digital quantification for spastic dysarthria detection and corticobulbar lesion localization.

Author summary

At the bedside, neurologists often detect dysarthria by listening to a patient’s speech, a practical but subjective approach that may delay diagnosis when dysarthria is mild. We developed a simple digital assessment approach that combines clinical phonetic knowledge with an accessible artificial intelligence tool, automatic speech recognition (ASR), to make the bedside judgement more objective while remaining clinically meaningful. Rather than relying on pre-existing open speech datasets, we newly collected speech recordings from patients with neuroimaging-confirmed corticobulbar lesions and matched healthy participants. Most patients had mild dysarthria, representing a setting where perceptual diagnosis is uncertain. We designed sentences that target different speech muscle groups and used ASR transcription errors as measurable features. A simple machine learning classifier was then trained on these features to evaluate diagnostic performance. The best sentence designs distinguished patients from controls with good performance and produced understandable results in relation to speech physiology. Our study illustrates a broader principle for digital neurology: artificial intelligence may be most useful when it is guided by clinical knowledge rather than replacing it and when it is available at the bedside. This approach could shift neurological assessment toward objective yet explainable way and could be extended beyond diagnosis to repeated monitoring during speech rehabilitation.

Article activity feed