Loading...
正在加载...
请稍候

[论文] Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree...

小凯 (C3P0) 2026年09月13日 00:47

论文概要

研究领域: cs.AI
作者: Bin Lei, Yu Li, Prafulla Kumar Choubey, Jiaxin Zhang, Becky Xiangyu Peng, Qinyuan Ye, Kartik Narayan, Caiwen Ding, Silvio Savarese, Chien-Sheng Wu
发布时间: 2026-09-13
arXiv: 2609.11061

中文摘要

树形结构化 rollout 为无评论家的可验证奖励强化学习(RLVR)提供步骤级信用:在中间点分叉链,兄弟结果差异估计步骤价值。每个分叉增加采样成本,因此实际预算通常每条链只允许少量分叉。在结果已基本确定的点放置分叉会产生大多同意的兄弟,几乎不提供信用信号;因此,对于给定树大小,分叉位置在很大程度上决定了步骤级 RL 能获得多少收益。大多数现有主流方法按结构放置分叉,如固定长度、中点和分隔符,或按下一个标记的熵。我们将分叉位置形式化为定位链价值曲线的枢轴,即预期结果转向的点。我们提出信念偏移分支:在候选边界读取模型的答案信念,并在连续信念分歧最大的步骤之前分叉。三种实例化跨越访问级别:黑盒探测、logit-lens 深度配置文件和学习的激活方向。在 RL 中,信念偏移分叉在数学聚合上领先,在 OLMo-3-7B 上比最强基线高 +2.6 聚合和 +2.9 AIME 2026,并在 OLMo 代码列上全面领先,LiveCodeBench-medium 高 +6.5。

原文摘要

Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step-level RL can gain. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next-token entropy. We formalize fork placement as locating the \emph{pivots} of the chain's value curve, where the expected outcome turns. We propose \emph{belief-shift branching}: read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most. Three instantiations, none needing step-level supervision, span access levels: a black-box probe, a logit-lens depth profile, and a learned activation direction, which is fit offline and therefore used only in the validation before RL training. The signal only \emph{places} forks, and the probe costs about \(1\%\) of step compute on mathematics and under \(5\%\) on code when it runs inside the rollout engine. In that validation, against Monte-Carlo value curves, a belief-shift signal ranks first in each of the eight model\(\times\)benchmark panels, ahead of entropy, structural, and LLM-judge baselines. In RL across three model families and two domains, belief-shift forking leads every mathematics aggregate, on OLMo-3-7B by \(+2.6\) aggregate and \(+2.9\) on AIME 2026 over the strongest baseline, and sweeps every OLMo code column, by \(+6.5\) on LiveCodeBench-medium.


自动采集于 2026-09-13

#论文 #arXiv #AI #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录