Paper Overview
Field: Computer Vision Authors: Mainak Singha, Niccolò Biondi, Elisa Ricci arXiv: 2508.03799 Code: https://github.com/mainaksingha01/DistMoE
Abstract
Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To address this, the authors propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning.
Key ideas
- Private FFN experts: In each layer of the language decoder, the public feedforward network (FFN) is augmented with a client-specific private FFN expert, whose goal is to acquire domain-specific knowledge.
- Scale mismatch problem: Independent expert training causes the private FFNs to learn representations of different scales and magnitudes, making it difficult to merge experts across clients.
- Public-anchored expert composition: To reduce client-specific drift, a composition stage updates only the router and lightweight private projection adapters on a mixture of local client data and public data, using an isotropic regularization loss. This enables rehearsal-free composition across clients—no raw data sharing or replay is required.
- Modular routing at inference: DistMoE routes tokens over both public and private experts, achieving token-level domain composition without explicit domain labels.
Results
Experiments across diverse visual-language benchmarks show that DistMoE enables flexible expert reuse, effective domain adaptation, and competitive performance while retaining modular control over client-specific knowledge.
---
*Auto-collected on 2026-08-12.*