A generalizable speech neuroprosthesis
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Intracortical brain-computer interfaces (BCIs) can restore communication to people with vocal tract paralysis by decoding cortical activity during attempted speech into text. State-of-the-art systems pairing neural-to-phoneme decoders with phoneme-to-word language models have achieved word error rates (WERs) as low as 1%, but only after collecting thousands of sentences of training data. Shortening the data collection process would facilitate scaling this new technology by reducing the time from device implant to high-accuracy communication. Here we introduce a transformer-based decoder model trained jointly across six intracortical speech BCI participants. For every participant — regardless of sex, disease etiology, or attempted speaking strategy — a multi-user model decoded speech more accurately (over 50% lower relative WER on average) than models trained on individual users’ data. Notably, the multi-user model could be finetuned on fewer than 200 sentences from a held-out user to achieve a WER below 7%. These results reveal how to pool intracortical data across people to yield more accurate, generalizable, and rapidly-deployable decoding models.