Paper Overview
Field: Machine Learning Authors: Maryam Maghsoudi, Shihab Shamma Published: 2025-05-07 arXiv: 2505.05131
Abstract
Decoding imagined speech from non-invasive brain recordings is challenging because imagined datasets are scarce and difficult to align temporally across subjects and sessions. In this work, the authors propose a new approach to decoding imagined speech that leverages the richer and more reliably labeled recordings made during listening to speech.
Method
Paired listened and imagined MEG recordings were collected from trained musicians exposed to rhythmic melodic and spoken stimuli. Using trained musicians helped improve temporal alignment across conditions.
A three-stage decoding pipeline was developed, revealing consistent and meaningful relationships between neural activity evoked by imagining versus listening to the same stimuli:
1. Imagined-to-listened mapping: Six linear and neural models were trained to map imagined MEG responses to listened responses. These models were evaluated zero-shot on unseen subjects to verify that the predicted listened responses preserve stimulus-specific information. 2. Word decoder: A contrastive word decoder was trained solely on listened MEG responses, evaluated with four embedding strategies covering semantic, acoustic, and phonetic representations. 3. Decoding imagined speech: Held-out subjects' imagined MEG responses were passed through the mapping pipeline to compute corresponding listened responses, which were then decoded by the listened-word decoder.
Results
Using rank-based analysis, the authors demonstrate that imagined words can be decoded at probabilities significantly above chance. All evaluations were performed on held-out subjects, making this a proof-of-concept for zero-shot imagined speech decoding. Performance improves as the amount of training data increases, indicating the approach is scalable and directly applicable to realistic brain-computer interface scenarios.
--- *Auto-collected on 2026-05-12*