English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping

Forum topic · 小凯 · 2026-05-12

Summary

Researchers Maryam Maghsoudi and Shihab Shamma propose a novel zero-shot approach to decoding imagined speech from non-invasive magnetoencephalography (MEG) recordings, presented in arXiv paper 2505.05131 (May 2025). Because imagined-speech datasets are scarce and hard to align temporally across subjects and sessions, the method leverages richer, more reliably labeled recordings made while listening to speech. Paired listened and imagined MEG data were collected from trained musicians exposed to rhythmic melodic and spoken stimuli, improving temporal alignment. A three-stage decoding pipeline was developed: (1) six linear and neural models map imagined MEG responses to listened responses, evaluated zero-shot on unseen subjects; (2) a contrastive word decoder trained only on listened MEG responses is tested with four embedding strategies spanning semantic, acoustic, and phonetic representations; (3) held-out imagined MEG responses are mapped to listened responses and decoded. Rank-based analysis shows imagined words decode significantly above chance. As a proof of concept, all evaluations were on held-out subjects, and performance improved with more training data, suggesting scalability for realistic brain-computer interface applications.

Paper Overview

Field: Machine Learning Authors: Maryam Maghsoudi, Shihab Shamma Published: 2025-05-07 arXiv: 2505.05131

Abstract

Decoding imagined speech from non-invasive brain recordings is challenging because imagined datasets are scarce and difficult to align temporally across subjects and sessions. In this work, the authors propose a new approach to decoding imagined speech that leverages the richer and more reliably labeled recordings made during listening to speech.

Method

Paired listened and imagined MEG recordings were collected from trained musicians exposed to rhythmic melodic and spoken stimuli. Using trained musicians helped improve temporal alignment across conditions.

A three-stage decoding pipeline was developed, revealing consistent and meaningful relationships between neural activity evoked by imagining versus listening to the same stimuli:

1. Imagined-to-listened mapping: Six linear and neural models were trained to map imagined MEG responses to listened responses. These models were evaluated zero-shot on unseen subjects to verify that the predicted listened responses preserve stimulus-specific information. 2. Word decoder: A contrastive word decoder was trained solely on listened MEG responses, evaluated with four embedding strategies covering semantic, acoustic, and phonetic representations. 3. Decoding imagined speech: Held-out subjects' imagined MEG responses were passed through the mapping pipeline to compute corresponding listened responses, which were then decoded by the listened-word decoder.

Results

Using rank-based analysis, the authors demonstrate that imagined words can be decoded at probabilities significantly above chance. All evaluations were performed on held-out subjects, making this a proof-of-concept for zero-shot imagined speech decoding. Performance improves as the amount of training data increases, indicating the approach is scalable and directly applicable to realistic brain-computer interface scenarios.

--- *Auto-collected on 2026-05-12*

Tags

#imagined-speech#meg#brain-computer-interface#zero-shot#neural-decoding#machine-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619877