Paper Overview
Field: Computer Vision Authors: Mainak Singha, Niccolò Biondi, Elisa Ricci Published: 2026-08-11 arXiv: 2508.03799
Abstract
Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To address this, the authors propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning.
Key Ideas
- Private FFN experts: In each layer of the language decoder, the public feedforward network (FFN) is augmented with a client-specific private FFN expert, aimed at acquiring domain-specific knowledge.
- Expert merging challenge: Independent expert training causes private FFNs to learn representations with different scales and magnitudes, making it difficult to merge experts.
- Public-anchored composition: To reduce client-specific drift, DistMoE introduces a public-anchored expert composition stage that updates only the router and lightweight private projection adapters on a mixture of local client data and public data, using an isotropic regularization loss. This enables rehearsal-free composition across clients without accessing raw private data from others.
- Token-level routing at inference: DistMoE performs modular routing over public and private experts, achieving token-level domain composition without requiring explicit domain labels.
- Paper: arXiv:2508.03799
- Code: https://github.com/mainaksingha01/DistMoE
Results
Experiments across diverse visual-language benchmarks show that DistMoE achieves flexible expert reuse, effective domain adaptation, and competitive performance, while preserving modular control over client-specific knowledge.