论文概要
研究领域: ML
作者: Semih Kara, Oğuzhan Ersoy
发布时间: 2026-06-09
arXiv: 2606.11173
中文摘要
自蒸馏通过匹配学生(仅见问题)和自教师(见问题+反馈上下文)的输出分布来训练模型。本文比较三种反馈条件:二值奖励(GRPO)、参考解、与求解器推理轨迹逐步对齐的批评。逐步对齐的批评效果最佳,超越GRPO 16.11分,超越参考解条件蒸馏5.27分。逐token优势分析揭示:对齐的反馈只针对推理失败的token,保留正确行为。
原文摘要
Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this improvement when the context is not present. The method works by matching the model's output distribution under two settings: a student that sees only the question, and a self-teacher that also sees the context. What the model learns therefore depends on what context the self-teacher receives, yet the design of this context remains largely unexplored. We study context design for self-distillation by training a solver on feedback from a frozen critic. We compare three conditions: (i) a binary reward (GRPO), (ii) the reference solution, and (iii) a step-by-step critique aligned to the solver's reasoning trace. Step-ali...
自动采集于 2026-06-11
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。