English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DistMoE: Rehearsal-Free Distributed Instruction Tuning with Private-Data MoE Routing

Forum topic · 小凯 · 2026-08-11

Summary

DistMoE is a mixture-of-experts (MoE) method for distributed visual instruction tuning of multimodal large language models (MLLMs), proposed by Mainak Singha, Niccolò Biondi, and Elisa Ricci (arXiv:2508.03799). Adapting MLLMs to diverse visual-language domains usually assumes centralized data access and costly joint training, which is restrictive when data is spread across private, domain-specific, or permission-limited clients. DistMoE augments the public feedforward network (FFN) in each language decoder layer with client-specific private FFN experts that capture domain-specific knowledge. Because independent expert training produces representations at inconsistent scales, the authors introduce a public-anchored expert composition stage: a router and lightweight private projection adapters are updated only on a mix of local client data and public data, with an isotropic regularization loss, enabling rehearsal-free composition across clients. At inference, DistMoE performs modular routing over public and private experts, achieving token-level domain composition without explicit domain labels. Experiments across diverse visual-language benchmarks show flexible expert reuse, effective domain adaptation, and competitive performance while preserving modular control over client-specific knowledge. Code is available at https://github.com/mainaksingha01/DistMoE.

Paper Overview

Field: Computer Vision Authors: Mainak Singha, Niccolò Biondi, Elisa Ricci arXiv: 2508.03799 Code: https://github.com/mainaksingha01/DistMoE

Abstract

Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To address this, the authors propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning.

Key ideas

  • Private FFN experts: In each layer of the language decoder, the public feedforward network (FFN) is augmented with a client-specific private FFN expert, whose goal is to acquire domain-specific knowledge.
  • Scale mismatch problem: Independent expert training causes the private FFNs to learn representations of different scales and magnitudes, making it difficult to merge experts across clients.
  • Public-anchored expert composition: To reduce client-specific drift, a composition stage updates only the router and lightweight private projection adapters on a mixture of local client data and public data, using an isotropic regularization loss. This enables rehearsal-free composition across clients—no raw data sharing or replay is required.
  • Modular routing at inference: DistMoE routes tokens over both public and private experts, achieving token-level domain composition without explicit domain labels.

Results

Experiments across diverse visual-language benchmarks show that DistMoE enables flexible expert reuse, effective domain adaptation, and competitive performance while retaining modular control over client-specific knowledge.

---

*Auto-collected on 2026-08-12.*

Tags

#mixture-of-experts#multimodal-llm#distributed-training#instruction-tuning#federated-learning#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633353