论文概要
研究领域: AI
作者: Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang
发布时间: 2026-08-25
arXiv: 2608.24870
中文摘要
组相对强化学习等待同一提示的兄弟推演,这对于长且可变的工具使用轨迹来说成本高昂。单流策略优化(SPO)通过持久提示级价值估计消除了这一依赖,但其方法在优化token均值actor损失之前对每个轨迹的优势进行白化。我们表明,轨迹中心化通常不会中心化actor消耗的token加权量,并通过在动作token测度下标准化终端结果优势来修复这一不匹配。此外,我们按生成提示证据的策略事件而非学习者接收顺序来组织提示证据。在ALFWorld两个模型规模和Math-TIR上的匹配运行中,SPO++提高了SPO的在线学习效率。成对消融实验确定动作token测度归一化是最强的测试组件。
原文摘要
Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level value estimate, but its recipe whitens one advantage per trajectory before optimizing a token-mean actor loss. We show that trajectory centering generally does not center the token-weighted quantity consumed by the actor, and fix the mismatch by standardizing terminal-outcome advantages under the action-token measure. We additionally organize prompt evidence by the policy event that generated it rather than learner receipt order. Across matched runs on ALFWorld at two model scales and on Math-TIR, SPO++ improves online learning efficiency over SPO. A pa...
自动采集于 2026-08-27
#论文 #arXiv #AI #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。