论文概要
研究领域: NLP
作者: Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao
发布时间: 2026-08-27
arXiv: 2608.27409
中文摘要
可验证奖励强化学习(RLVR)改善大语言模型的特定能力,但覆盖多种能力通常涉及训练单独的领域专家然后整合它们。我们按复用的产物组织三种融合范式:Merge组合专家任务向量,Mix RL汇集它们的数据集,多教师在线策略蒸馏(MOPD)同时使用两者。由于它们大多被孤立研究,如何比较和如何选择仍不清楚。我们使用共享专家和数据在模型规模和多领域基准套件上比较三者。尽管平均性能最多相差1.4分,但单一基准上差距达8.6分。训练动态暴露了不同约束:Mix RL依赖领域混合比例,MOPD受教师限制,Merge将所有专家更新压缩为一个。所有三种都提高了单样本准确率,但在解决方案覆盖或保留能力方面没有可测量的增益或损失。实用指南:专家已存在且廉价融合最重要时用Merge;没有专家训练统一模型用Mix RL;保留领域特定增益比超越教师更重要时用MOPD。
原文摘要
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artefacts they reuse: Merge combines expert task vectors, Mix RL pools their datasets, and multi-teacher on-policy distillation (MOPD) uses both. Because they have largely been studied in isolation, how they compare and how to choose among them remain unclear. We compare all three using shared experts and data across model scales and a multi-domain benchmark suite. Although their average performance differs by at most 1.4 points, the gap reaches 8.6 points on a single benchmark, with domain-level variation tracking cross-...
自动采集于 2026-08-30
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。