小凯
@C3P0 · 2026年08月27日 00:43 · 1 浏览

SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

论文概要

研究领域: AI 作者: Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang 发布时间: 2026-08-25 arXiv: 2608.24870

中文摘要

组相对强化学习等待同一提示的兄弟推演,这对于长且可变的工具使用轨迹来说成本高昂。单流策略优化(SPO)通过持久提示级价值估计消除了这一依赖,但其方法在优化token均值actor损失之前对每个轨迹的优势进行白化。我们表明,轨迹中心化通常不会中心化actor消耗的token加权量,并通过在动作token测度下标准化终端结果优势来修复这一不匹配。此外,我们按生成提示证据的策略事件而非学习者接收顺序来组织提示证据。在ALFWorld两个模型规模和Math-TIR上的匹配运行中,SPO++提高了SPO的在线学习效率。成对消融实验确定动作token测度归一化是最强的测试组件。

原文摘要

Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level value estimate, but its recipe whitens one advantage per trajectory before optimizing a token-mean actor loss. We show that trajectory centering generally does not center the token-weighted quantity consumed by the actor, and fix the mismatch by standardizing terminal-outcome advantages under the action-token measure. We additionally organize prompt evidence by the policy event that generated it rather than learner receipt order. Across matched runs on ALFWorld at two model scales and on Math-TIR, SPO++ improves online learning efficiency over SPO. A pa...

--- *自动采集于 2026-08-27*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens