[论文] Decoupling Exploration from Optimization in RLVR
研究领域: NLP 作者: Saif Punjwani, Micah Goldblum 发布时间: 2026-10-07 arXiv: 2610.10536
论文概要
研究领域: NLP 作者: Saif Punjwani, Micah Goldblum 发布时间: 2026-10-07 arXiv: 2610.10536
中文摘要
现代语言模型在已训练好的检查点之上进行基于可验证奖励的强化学习(RLVR)。RLVR的一个关键承诺是发现新的推理策略——原则上,模型可以采样出训练数据中不存在的创新思路。然而实践中,为RLVR添加强力的新颖性激励收效有限,甚至会降低模型质量。由于可验证奖励只监督模型知识和行为的狭窄切面,这种退化很难恢复。为此,我们将探索与优化解耦,提出Exploration-Distillation(ExpDis)框架。我们用一个或多个带有新颖性奖励的探索策略进行训练,筛选其中正确且高质量的轨迹,将其蒸馏到独立的学生策略中。学生策略随后在不带新颖性奖励的情况下训练。上述过程重复多轮,在探索与优化之间交替。这种解耦使我们能够激进地扩展探索规模而不损害学生策略。在七个数学推理基准和两个模型家族上,ExpDis在相同墙钟预算下优于DAPO。此外,我们观察到pass@k扩展性改善,表明ExpDis产生的模型能生成更多样化的正确解。
原文摘要
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them ...
*自动采集于 2026-10-09*
#论文 #arXiv #NLP #小凯