论文概要
研究领域: ML
作者: Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich
发布时间: 2026-09-25
arXiv: 2609.27035
中文摘要
GRPO 及相关策略梯度方法训练语言模型智能体时,把整个多轮 rollout 压缩成单一标量轨迹奖励再进入策略更新。当任务由不同技能组成、尤其在稀疏且延迟的环境反馈下,这种压缩是有损的:优化器必须隐式推断哪种能力驱动了结果,以及应如何改变行为。我们认为正确的原语不是更好的标量,而是分解:轨迹奖励应在进入策略更新前按子任务拆分。我们提出 RLDS,核心是子任务分解优势估计(SDAE):替代 GRPO 标量优势,将轨迹奖励按固定分类法拆分为各子任务份额,为每个子任务计算组相对优势,并按各子任务重要性加权把 token 级信用集中到反思标记该子任务执行具关键意义的步骤附近。我们在四个智能体基准上评估:FrozenLake(稀疏网格导航)、HotpotQA(多跳问答,单检索工具)、ScienceWorld(长程具身科学)、DeepResearch(长文研究,四工具,复合评分奖励)。训练期间输出的异质性诊断显示分解何时见效:收益随子任务异质性扩大,在高异质性任务 ScienceWorld(+11.5 分,配对自举 95% CI [+9.8, +13.3])与 FrozenLake(+9.8 分,[+7.0, +12.8])上最大,在 HotpotQA 与 DeepResearch 上与噪声无异——诊断早已预测这两处无利可图。ScienceWorld 在 RLDS 下还比标量 GRPO 更省算力(每步墙钟时间 -10.9%),因为长 rollout 摊薄了固定的反思-评分开销。
原文摘要
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a better scalar but a decomposition: trajectory reward should be split along subtasks before it enters the policy update. We introduce Reinforcement Learning with Decomposed Subtasks (RLDS), whose core is Subtask-Decomposed Advantage Estimation (SDAE): a replacement for the scalar GRPO advantage that sp...
自动采集于 2026-09-25
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。