Loading...
正在加载...
请稍候

[论文] Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text

小凯 (C3P0) • 2026年10月02日 00:42

论文概要

研究领域: ML
作者: Dulhan Jayalath, Oiwi Parker Jones
发布时间: 2026-09-30
arXiv: 2609.26742

中文摘要

我们发现,从非侵入式脑电记录中解码文字所报告的重大改进,在很大程度上可以在完全不用脑数据的情况下复现。在 d'Ascoli 等人(2025)的有影响力的工作中,来自感知连续语音的受试者脑活动时间序列被分割为从每个词开始的固定长度窗口。神经网络随后联合生成句子中所有词的预测。相邻窗口部分重叠,隐含地揭示了词与词之间的间隔。由于这些间隔指示了所说词的时长,而不同词往往有不同的时长——例如 "the" 比 "supercalifragilisticexpialidocious" 短得多——神经网络可以在不依赖底层脑活动的情况下改进其词预测。与此一致的是,该方法在不含任何脑信息的合成信号上达到了 22.0% 的平衡准确率,相比之下真实脑电记录上的结果为 22.3%。为防止网络学习这一捷径,我们做了一个简单而关键的改动:不再联合编码句子中的所有窗口,而是独立处理每个窗口。结果是,神经网络通过从脑电记录中学习词特定信息而取得了更好的性能。这使两种现有策略变得比以前有效得多:聚合来自同一词不同神经响应的预测,以及使用预训练 LLM 作为语言先验,现在都大幅改进了结果。在我们的感知语音基准上,这一简单方法(SimpleB2T)在每次五个观测的条件下实现了 36.6% 的词错误率,接近过去侵入式语音解码的性能(尽管条件不同)。本工作揭示了脑到文字解码中的一个重要捷径,并表明移除它会产生一种简单且显著更有效的策略。

原文摘要

We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length windows starting at each word. A neural network then generates predictions for all of the words in a sentence together. Neighbouring windows partially overlap, implicitly revealing the interval between words. Since these intervals indicate the duration of the words spoken, and different words tend to have different durations - for example, 'the' is much shorter than 'supercalifragilisticexpialidocious' - the neural network can improve its predictions of words without relying on the underlying bra...


自动采集于 2026-10-02

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录