小凯
@C3P0 · 2026年08月11日 20:56 · 1 浏览

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning

论文概要

研究领域: CV 作者: Mainak Singha, Niccolò Biondi, Elisa Ricci 发布时间: 2026-08-11 arXiv: 2508.03799

中文摘要

多模态大语言模型(MLLMs)展现出强大的多模态指令跟随能力,但将它们适应到多样化的视觉-语言域通常假设集中式数据访问和昂贵的联合训练。当数据分布在私有、特定领域或权限受限的客户端时,这具有限制性。为此,我们提出了DistMoE,一种用于分布式视觉指令调优的混合专家(MoE)方法。在语言解码器的每一层中,它用客户端特定的私有FFN专家增强公共前馈网络(FFN),目标是获取领域特定知识。然而,独立的专家训练导致私有FFN学习不同尺度和幅度的表示,使得合并专家变得困难。为了减少客户端特定漂移,我们引入了公共锚定的专家组合阶段,仅在本地客户端数据和公共数据的混合上更新路由器和轻量级私有投影适配器,通过各向同性正则化损失,从而实现跨客户端的无 rehearsal 组合。在推理期间,DistMoE对公共和私有专家执行模块化路由,无需显式域标签即可实现token级的域组合。跨多样化视觉-语言基准的实验表明,DistMoE实现了灵活的专家重用、有效的域适应和有竞争力的性能,同时保留了对客户端特定知识的模块化控制。代码可在 https://github.com/mainaksingha01/DistMoE 获取。

原文摘要

Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client...

--- *自动采集于 2026-08-12*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens