Overview
Field: Computer Vision Authors: Mainak Singha, Niccolò Biondi, Elisa Ricci Published: 2026-08-11 arXiv: 2508.03799
Key Points
- Multimodal Large Language Models (MLLMs) show strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training — restrictive when data is distributed across private, domain-specific, or permission-limited clients.
- DistMoE is a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder, it augments the public feedforward network (FFN) with a client-specific private FFN expert to acquire domain-specific knowledge.
- Independent expert training causes private FFNs to learn representations at different scales and magnitudes, making expert merging difficult.
- To reduce client-specific drift, the authors introduce a public-anchored expert composition stage that updates only the router and lightweight private projection adapters on a mixture of local client data and public data, using an isotropic regularization loss — enabling rehearsal-free composition across clients.
- During inference, DistMoE performs modular routing over public and private experts, achieving token-level domain composition without explicit domain labels.
- Experiments across diverse visual-language benchmarks show flexible expert reuse, effective domain adaptation, and competitive performance while retaining modular control over client-specific knowledge.
- Paper: https://arxiv.org/abs/2508.03799
- Code: https://github.com/mainaksingha01/DistMoE