论文概要
研究领域: NLP
作者: Junshu Pan, Zhizhang Fu, Shulin Huang, et al.
发布时间: 2026-09-30
arXiv: 2609.26735
中文摘要
基于可验证奖励的强化学习(RLVR)提升了大语言模型的推理能力,但其预测仍对任务无关的提示词特征敏感。我们通过半事实(semifactual)提示词干预来研究这种敏感性——这些干预保持了底层问题及其答案不变。分析揭示了 token 级敏感性的显著差异,并表明在解码过程中抑制高漂移 token 候选可以在不更新模型权重的情况下提高推理准确率。这些发现凸显了 GRPO(组相对策略优化)的一个局限性:它给每个响应 token 分配相同的结果衍生优势,可能在强化有用推理的同时,也强化了潜在的虚假依赖。受此启发,我们引入半事实信用增强策略优化(SCAPO),这是一个受因果启发的 GRPO 变体,将半事实稳定性纳入 token 级信用分配。SCAPO 在半事实干预下测量固定响应的 token 概率漂移,并使用归一化稳定性分数,在训练早期降低相对不稳定 token 的优势,但仅凭稳定性不给予额外信用。在 Qwen3-4B-Base 和 Qwen3-1.7B-Base 上,SCAPO 相比 GRPO 分别将 AIME 2024-2026 准确率提升了 5.63 和 4.17 个百分点。在两个模型规模上,SCAPO 在大多数数学基准测试和所有分布外基准测试中均取得了比较方法中的最佳结果。
原文摘要
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we intr...
自动采集于 2026-10-02
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。