Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
Abstract
We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length windows starting at each word. A neural network then generates predictions for all of the words in a sentence together. Neighbouring windows partially overlap, implicitly revealing the interval between words. Since these intervals indicate the duration of the words spoken, and different words tend to have different durations - for example, "the" is much shorter than "supercalifragilisticexpialidocious" - the neural network can improve its predictions of words without relying on the underlying brain activity. Consistent with this, the method reaches 22.0% balanced accuracy on synthetic signals containing no brain information, compared with 22.3% on real brain recordings. To prevent the network from learning this shortcut, we make a single, simple change. Instead of jointly encoding all windows in a sentence, we process each independently. As a result, the neural network achieves better performance by learning underlying word-specific information from brain recordings. This makes two existing strategies become much more effective than before. Both aggregating predictions from distinct neural responses to the same word and using a pretrained LLM as a linguistic prior now substantially improve results. On our perceived speech benchmark, this simple recipe (SimpleB2T) achieves a word error rate of 36.6% with five observations per word, approaching past invasive speech decoding performance, albeit under different conditions. The results in this work expose an important shortcut in brain-to-text decoding and show that removing it leads to a simple and considerably more effective strategy.
Community
We show that a widely used non-invasive brain-to-text setup can get much of its apparent performance from a timing shortcut. Overlapping word-aligned brain windows reveal how long words last, even without useful brain information. Removing this shortcut changes our view of progress in the field and leads to a simpler decoder that combines information decoded from brain activity with repeated observations and an LLM, reaching 36.6% WER on a clinically motivated communication benchmark.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- HDND: Hierarchical Dynamic Neural Decoding for Multilingual Word/Character Retrieval from Non-Invasive Brain Recordings (2026)
- Integrating Language Models into Listened and Imagined Speech Decoding from MEG (2026)
- The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding (2026)
- Decoding silent reading from non-invasive EEG (2026)
- BAT-CLIP: Trimodal Alignment of Brain, Audio and Text (2026)
- SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding (2026)
- Brain2Speech-Net: Fast and Intelligible Brain-to-Speech Synthesis Without Text Decoding (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.40359 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper