Paper Overview
- Field: Machine Learning
- Authors: Semih Kara, Oğuzhan Ersoy
- Published: 2026-06-09
- arXiv: 2606.11173
- Self-distillation quality depends critically on what context the self-teacher receives, an underexplored design choice.
- Three feedback conditions compared: binary reward (GRPO), reference solution, and step-aligned critique from a frozen critic.
- Step-aligned critiques win: +16.11 points over GRPO, +5.27 points over reference-solution distillation.
- Token-level analysis shows aligned feedback selectively corrects failing tokens while preserving already-correct behavior.
Abstract
Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this improvement when the context is not present. The method works by matching the model's output distribution under two settings: a student that sees only the question, and a self-teacher that also sees the context. What the model learns therefore depends on what context the self-teacher receives, yet the design of this context remains largely unexplored. We study context design for self-distillation by training a solver on feedback from a frozen critic. We compare three conditions: (i) a binary reward (GRPO), (ii) the reference solution, and (iii) a step-by-step critique aligned to the solver's reasoning trace. Step-aligned critiques work best, outperforming GRPO by 16.11 points and the reference solution by 5.27 points. Token-level advantage analysis reveals that aligned feedback targets only tokens where reasoning failed, preserving correct behavior.
Key Takeaways
*Auto-collected on 2026-06-11.*