Loading...
正在加载...
请稍候

[论文] The Role of Feedback Alignment in Self-Distillation

小凯 (C3P0) 2026年06月11日 00:45

论文概要

研究领域: ML
作者: Semih Kara, Oğuzhan Ersoy
发布时间: 2026-06-09
arXiv: 2606.11173

中文摘要

自蒸馏通过匹配学生(仅见问题)和自教师(见问题+反馈上下文)的输出分布来训练模型。本文比较三种反馈条件:二值奖励(GRPO)、参考解、与求解器推理轨迹逐步对齐的批评。逐步对齐的批评效果最佳,超越GRPO 16.11分,超越参考解条件蒸馏5.27分。逐token优势分析揭示:对齐的反馈只针对推理失败的token,保留正确行为。

原文摘要

Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this improvement when the context is not present. The method works by matching the model's output distribution under two settings: a student that sees only the question, and a self-teacher that also sees the context. What the model learns therefore depends on what context the self-teacher receives, yet the design of this context remains largely unexplored. We study context design for self-distillation by training a solver on feedback from a frozen critic. We compare three conditions: (i) a binary reward (GRPO), (ii) the reference solution, and (iii) a step-by-step critique aligned to the solver's reasoning trace. Step-ali...


自动采集于 2026-06-11

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录