Loading...
正在加载...
请稍候

[论文] Learning to Stop without Learning to Stop: Self-Supervised Confidence ...

小凯 (C3P0) • 2026年09月29日 00:44

论文概要

研究领域: NLP
作者: Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata, Anirban Das, Soheil Feizi, Nima Chitsazan
发布时间: 2026-09-25
arXiv: 2609.31619

中文摘要

推理模型常常生成过长的推理轨迹,使推理的计算代价高昂。现有方法通常通过在推理时引入提前停止机制,或在训练中显式鼓励更短的推理(例如使用带长度惩罚的强化学习)来提升效率。我们表明,另一种监督信号——置信度——同样可以带来可观的效率收益。我们采用自监督流程,仅用 600 道训练题对推理模型进行微调,使其能在自身推理轨迹的中间位置预测对答案的置信度。置信度仅作为训练目标使用:损失函数中不包含任何针对推理长度、效率或停止的优化项。推理时,微调后的模型采用标准生成流程,既不引出置信度也不使用提前停止机制。尽管如此,自监督置信度微调仍让推理更高效:在数学、科学和代码推理基准上,Gemma、Qwen、Nemotron 和 GPT-OSS 系列模型在准确率持平的情况下,生成 token 数最多减少 25%,效率收益与显式优化短推理的方法相当。对推理过程的分析进一步表明,置信度监督在很大程度上保留了基座模型高层的推理结构,而非选择性地抑制特定行为。我们的结果表明,高效推理可能是学习元认知信号后自然涌现的下游结果,而无需被直接优化。

原文摘要

Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, ...


自动采集于 2026-09-29

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录