English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MOJO: Joint Self-Supervised and Supervised Training for Generalizable Neural Population Decoding

Forum topic · 小凯 · 2026-07-17

Summary

Researchers introduce MOJO (Masked autOencoder-based JOint training), a framework for spike-tokenizing neural models that combines self-supervised learning (SSL) via masked autoencoding with supervised learning (SL). Current spike-based neural decoders rely solely on SL, restricting training to datasets with paired behavioral labels. MOJO removes this limitation by jointly leveraging unlabelled neural data. Evaluated on three spiking datasets—monkey motor cortex recordings during reaching tasks and multi-regional mouse visual/decision-making recordings—MOJO outperforms purely supervised models, especially with limited labelled data and in few-shot fine-tuning on new sessions. SSL also yields more interpretable neuron representations, improving brain-region classification and spike-statistics prediction without explicit optimization. The approach generalizes beyond spikes to human speech-cortex electrocorticography, matching neural foundation models designed for continuous signals. Results point toward more flexible, scalable data usage for training neural foundation models across tasks, species, and modalities.

Paper Overview

Field: Machine Learning Authors: Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo, Matthew G. Perich, Guillaume Lajoie Published: 2026-07-15 arXiv: 2607.14086

Summary

Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning (SL), limiting training to datasets with paired behavioural labels.

To address this limitation, the authors introduce MOJO (Masked autOencoder-based JOint training), a training framework for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives.

Key Findings

  • Evaluated on three spiking datasets, spanning monkey motor cortex activity during reaching tasks and multi-regional mouse recordings covering visual and decision-making tasks.
  • MOJO outperforms purely SL-trained models, with the improvement being especially significant when trained with limited labelled data, and in few-shot fine-tuning where only small amounts of labelled data are available for new sessions.
  • Incorporating SSL produces more interpretable neuron representations, improving brain-region classification and spike-statistics prediction without explicitly optimizing for these tasks.
  • MOJO generalizes beyond spike data to human speech-cortex electrocorticography (ECoG), continuing to outperform pure SL models and achieving performance comparable to neural foundation models (NFMs) specifically designed for continuous signals.

Implications

Overall, augmenting spike-tokenized models with SSL improves performance in label-scarce settings, enables the use of unlabelled data across diverse tasks and species, and generalizes to other neural modalities. These results point the way toward more flexible and scalable data usage when training neural foundation models.

--- *Source: arXiv:2607.14086*

Tags

#machine-learning#neural-decoding#self-supervised-learning#brain-computer-interface#neuroscience#arxiv#masked-autoencoder#neural-foundation-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178395196