论文概要
研究领域: cs.AI
作者: Bin Lei, Yu Li, Prafulla Kumar Choubey, Jiaxin Zhang, Becky Xiangyu Peng, Qinyuan Ye, Kartik Narayan, Caiwen Ding, Silvio Savarese, Chien-Sheng Wu
发布时间: 2026-09-13
arXiv: 2609.11061
中文摘要
树形结构化 rollout 为无评论家的可验证奖励强化学习(RLVR)提供步骤级信用:在中间点分叉链,兄弟结果差异估计步骤价值。每个分叉增加采样成本,因此实际预算通常每条链只允许少量分叉。在结果已基本确定的点放置分叉会产生大多同意的兄弟,几乎不提供信用信号;因此,对于给定树大小,分叉位置在很大程度上决定了步骤级 RL 能获得多少收益。大多数现有主流方法按结构放置分叉,如固定长度、中点和分隔符,或按下一个标记的熵。我们将分叉位置形式化为定位链价值曲线的枢轴,即预期结果转向的点。我们提出信念偏移分支:在候选边界读取模型的答案信念,并在连续信念分歧最大的步骤之前分叉。三种实例化跨越访问级别:黑盒探测、logit-lens 深度配置文件和学习的激活方向。在 RL 中,信念偏移分叉在数学聚合上领先,在 OLMo-3-7B 上比最强基线高 +2.6 聚合和 +2.9 AIME 2026,并在 OLMo 代码列上全面领先,LiveCodeBench-medium 高 +6.5。
原文摘要
Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step-level RL can gain. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next-token entropy. We formalize fork placement as locating the \emph{pivots} of the chain's value curve, where the expected outcome turns. We propose \emph{belief-shift branching}: read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most. Three instantiations, none needing step-level supervision, span access levels: a black-box probe, a logit-lens depth profile, and a learned activation direction, which is fit offline and therefore used only in the validation before RL training. The signal only \emph{places} forks, and the probe costs about \(1\%\) of step compute on mathematics and under \(5\%\) on code when it runs inside the rollout engine. In that validation, against Monte-Carlo value curves, a belief-shift signal ranks first in each of the eight model\(\times\)benchmark panels, ahead of entropy, structural, and LLM-judge baselines. In RL across three model families and two domains, belief-shift forking leads every mathematics aggregate, on OLMo-3-7B by \(+2.6\) aggregate and \(+2.9\) on AIME 2026 over the strongest baseline, and sweeps every OLMo code column, by \(+6.5\) on LiveCodeBench-medium.
自动采集于 2026-09-13
#论文 #arXiv #AI #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。