论文概要
研究领域: CV
作者: Mainak Singha, Niccolò Biondi, Elisa Ricci
发布时间: 2026-08-12
arXiv: 2508.05146
中文摘要
多模态大语言模型(MLLM)展现出强大的多模态指令跟随能力,但将它们适配到多样化的视觉-语言域通常假设可以集中访问数据并进行代价高昂的联合训练。当数据分布在私有、特定域或权限受限的客户端时,这一假设就具有很大局限性。为此,我们提出DistMoE,一种用于分布式视觉指令微调的混合专家(MoE)方法。在语言解码器的每一层中,它在公共前馈网络(FFN)基础上增加一个客户端特定的私有FFN专家,以获取域特定知识。然而,独立专家训练导致私有FFN学习到不同尺度和量级的表示,使得专家合并变得困难。为减少客户端特定漂移,我们引入公共锚定专家组合阶段,仅在本地客户端数据和公共数据的混合上更新路由器和轻量级私有投影适配器,通过各向同性正则化损失实现,从而使其成为跨客户端免排练组合。推理时,DistMoE在公共和私有专家上执行模块化路由,实现无需显式域标签的逐token域组合。在多样化视觉-语言基准上的实验表明,DistMoE实现了灵活的专家重用、有效的域适配和具有竞争力的性能,同时保留了对客户端特定知识的模块化控制。
原文摘要
Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client...
自动采集于 2026-08-12
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。