Limits of Trial-Adaptive Neural–Language Fusion Across Large Language Models in P300 Brain–Computer Interfaces
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Objective
Published P300-speller fusion schemes fix prior trust regardless of trial reliability; we tested whether a reliability estimate improves on it.
Methods
We reanalyzed 3,373 archived P300-speller selections from 47 people with ALS (BigP3BCI). A fair, matched-search-space comparison, tuning both a fixed weight and an adaptive policy out-of-fold, was evaluated across 22 evaluable language-model priors up to 46.7B parameters. Two representative priors, GPT-2 and a classical 5-gram, additionally received detailed naive and mechanistic analyses.
Results
No prior’s 95% CI favored adaptive fusion under the fair comparison, despite unexploited oracle headroom at every scale. Under GPT-2, the naive comparison was significantly worse for adaptive fusion; both anchors converged to a degenerate or near-degenerate fair-comparison solution. For the representative anchors, three further controllers failed to convert that headroom into benefit; the fixed-fused posterior’s output probability outperformed the best controller for flagging errors (2.8- to 3.8-fold enrichment).
Conclusion
A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage.
Significance
Adaptive weighting should be validated against a fairly tuned baseline across model families and scales; in this dataset, the fused output’s confidence identified high-risk selections better than the tested purpose-built ranker.