[论文] Planning to Learn

研究领域: ML 作者: Ian Osband 发布时间: 2026-10-02 arXiv: 2610.03667

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Ian Osband 发布时间: 2026-10-02 arXiv: 2610.03667

中文摘要

策略梯度方法是现代强化学习的核心,包括 LLM 后训练。当它们表现不佳时,通常的怀疑对象是探索、信用分配和动作采样噪声。而分类问题没有这些困难。分类器是一个策略,其期望奖励——期望准确率——就是它对正确标签赋予的概率,且由于标签已知,策略梯度是精确且光滑的。然而,精确策略梯度在期望准确率上却输给了交叉熵。精确梯度是短视的:它只根据当前收益来评价一次更新,但每次更新同时决定了下一次更新的起点,所以一次更新的价值取决于还剩多少学习空间。从这个角度看,交叉熵是'有耐心的准确率'——一个样本在对数几率以单位速度永远上升的情况下需要支付的总误差,而精确策略梯度是零时域极限。在剩余学习量处截断这个总量就得到 horizon loss——一行代码的改动,随着训练耗尽从交叉熵移向精确策略梯度。在一个简单的分配模型中,它可证明地逃脱了困住两个端点的陷阱。在 MNIST 和 ImageNet(ResNet-50、ResNet-101 和 ViT-S/16)上,horizon loss 在平坦学习率下提升了 top-1 准确率,且增益随标签噪声增大而增长。

原文摘要

Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update's value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy, the total error an example would pay if its lo...


*自动采集于 2026-10-06*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens