English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DistMoE: Rehearsal-Free Expert Routing for Distributed Multimodal Instruction Tuning

Forum topic · 小凯 · 2026-08-12

Summary

DistMoE (arXiv:2508.05146) is a mixture-of-experts method for distributed visual instruction tuning of multimodal large language models without centralized data access. Each language decoder layer augments a public feedforward network (FFN) with a client-specific private FFN expert that captures domain-specific knowledge. Because independent expert training produces private FFNs with differing representation scales, making expert merging difficult, the authors introduce a public-anchored expert composition stage: the router and lightweight private projection adapters are updated only on a local mix of client and public data under an isotropic regularization loss, enabling rehearsal-free composition across clients. At inference, DistMoE routes tokens modularly between public and private experts, achieving per-token domain composition without explicit domain labels. Experiments on diverse visual-language benchmarks show flexible expert reuse, effective domain adaptation, competitive performance, and modular control of client-specific knowledge. Suitable for federated or privacy-constrained multimodal adaptation scenarios.

Overview

Field: Computer Vision Authors: Mainak Singha, Niccolò Biondi, Elisa Ricci Published: 2026-08-12 arXiv: 2508.05146

Abstract

Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients.

To this end, the authors propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal of acquiring domain-specific knowledge. However, independent expert training causes the private FFNs to learn representations of different scales and magnitudes, making merging the experts difficult.

To reduce client-specific drift, DistMoE introduces a public-anchored expert composition stage: the router and lightweight private projection adapters are updated only on a local mix of client data and public data, via an isotropic regularization loss, making it a rehearsal-free composition across clients. At inference time, DistMoE performs modular routing over public and private experts, enabling per-token domain composition without explicit domain labels.

Experiments on diverse visual-language benchmarks show that DistMoE achieves flexible expert reuse, effective domain adaptation, and competitive performance while retaining modular control over client-specific knowledge.

Key points

  • Problem: Adapting MLLMs across domains usually requires centralized data and joint training; impractical when data is private or distributed across clients.
  • Approach: Per-layer private FFN experts added to a public FFN backbone for distributed visual instruction tuning.
  • Challenge addressed: Independent expert training causes scale/magnitude drift across private FFNs, hindering expert merging.
  • Solution: Public-anchored composition stage with isotropic regularization, updating only router and lightweight projection adapters on local client+public data — no data rehearsal required.
  • Inference: Modular token-level routing across public/private experts without explicit domain labels.
  • Results: Flexible expert reuse, effective domain adaptation, and competitive performance on diverse visual-language benchmarks.

Tags

#mixture-of-experts#multimodal-llm#federated-learning#visual-instruction-tuning#domain-adaptation#paper#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633375