Intent Drift in LLM-Assisted BCI Communication: An In-Silico Benchmark Under Simulated Decoder Corruption

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Large language models (LLMs) are increasingly used to correct noisy, error-prone text typed through brain-computer interfaces (BCIs) and other assistive communication devices. A fluent correction can still be wrong: the model can substitute a different intended message, a failure we call intent drift. Whether meaning survives correction, and whether model confidence flags failure, is unmeasured.

Methods

We built an in-silico benchmark: 20 open-weight LLMs corrected simulated P300 speller text under five levels of decoder error (0-40% character error rate). Testing used the full ALS message-banking vocabulary, a set of clinically important messages (eg, involving a medication dose), and matched controls. Each of 4,252,326 outputs was scored faithful, degraded, or drifted by an automated system benchmarked against physicians. A separate subanalysis compared six correction strategies across seven models.

Findings

Drift rose steeply with decoder error, from 2.2% at no error to 60.3% at the most severe level tested (odds ratio 2.30 per 10-percentage-point increase)-a stress-test ceiling, not an expected real-world rate. Model confidence distinguished correct from incorrect outputs reasonably well (AUROC 0.83) but overstated its own reliability: 28.4% of high-confidence outputs were not faithful. Clinically important messages drifted slightly more than matched controls (odds ratio 1.10). At low error rates, corrections succeeded more often than they went confidently wrong, but this reversed above 20-30% error. No correction strategy eliminated drift: cautious approaches reduced it, permissive ones increased it, and even the best still produced a wrong message in 18 of 100 corrections. Physician review of a sample agreed moderately with automated scoring (kappa 0.41); adjusting for this disagreement lowered but did not remove the pattern of rising drift with decoder error (31.4% to 28.3%).

Interpretation

LLMs correcting BCI text produced fluent but sometimes wrong messages, more often as decoding quality worsened, and their own confidence did not reliably warn when this happened. These in-silico findings support evaluating such systems for meaning preservation, not only speed and accuracy, before deployment; prospective, human-in-the-loop evaluation is needed.

Funding

Harvard Catalyst (CTSA UL1TR002541).

Research in context

Evidence before this study

We searched PubMed, arXiv, and bioRxiv from database inception to July 19, 2026, combining terms for large language models with brain-computer interfaces, P300 spellers, and augmentative and alternative communication, without language restriction. Exact strings and record counts are in eMethods S9. Prior studies integrating large language models with P300 spellers and communication BCIs reported keystroke savings, typing speed, information transfer rate, and character-level accuracy. Some concluded that correction was near-optimal once residual errors were manually fixed, placing meaning-level change outside their frame. No study had measured how often post-editing changed a message’s intent, whether this depended on decoder corruption, or whether stated confidence tracked whether the output was faithful.

Added value of this study

This in-silico benchmark measures intent drift and confidence calibration in language-model post-editing as a function of decoder corruption, reported separately across AUTH, a message-critical set, and matched controls. Using an empirical P300 confusion matrix and 4,252,326 labeled generations across 20 open-weight models, the primary benchmark questions found detected drift rose steeply with corruption and stated confidence discriminated faithful outputs well but was poorly calibrated; secondary analyses found message-critical content carried excess drift that survived detector removal and that detected drift varied more than two-fold across models (20.8-49.3% in AUTH). A separate subanalysis of six correction strategies (original seven-model panel) found no strategy removed drift: conservative editing or abstention lowered it, while alternatives or expansion raised it.

Implications of all the available evidence

Evaluations of language-model-assisted communication BCIs and AAC systems should report intent drift and confidence calibration alongside speed and accuracy, should test message-critical content separately from routine messages, and should treat interface policy as a measured design variable. Because the construct is communicative intent, the absence of patient and public involvement in probe-set design is material. Because confidence did not reliably flag drift, it alone cannot gate human review. These are in-silico findings; they support the need for prospective, human-in-the-loop evaluation before deployment implications can be drawn.

Article activity feed