English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Dense Supervision, Sparse Updates: The Sparsity and Geometry of On-Policy Distillation

Forum topic · 小凯 · 2026-06-14

Summary

This paper analyzes on-policy distillation (OPD), a training method that combines on-policy student trajectories with dense teacher supervision. The authors find that OPD-style parameter updates are small in magnitude and coordinate-sparse, yet distributed across model layers and typically concentrated in the FFN (feed-forward network) components. Although the updates are numerically full-rank, their spectrum is concentrated, indicating low effective dimensionality. These findings suggest that dense teacher supervision does not transform OPD into ordinary dense parameter rewriting; instead, OPD retains important geometric signatures characteristic of on-policy post-training. The work provides an empirical characterization of how distillation updates are structured in large language models. Paper: arXiv:2606.13657.

Paper Overview

  • Field: Machine Learning
  • Authors: Guo Yu, Wenlin Liu, Yulan Hu, Hao-Xuan Ma, Jun-Peng Jiang, Han-Jia Ye
  • Published: 2026-06-11
  • arXiv: 2606.13657
  • Abstract

    On-policy distillation (OPD) combines on-policy student trajectories and dense teacher supervision. Our analysis shows OPD-style updates are small and coordinate-sparse, distributed across layers and usually FFN-heavy. The updates are numerically full-rank but spectrally concentrated. These findings suggest dense teacher supervision does not turn OPD into ordinary dense parameter rewriting; OPD retains important geometric signatures of on-policy post-training.

    Key Findings

  • Sparse updates: OPD parameter updates are small in magnitude and coordinate-sparse, despite the dense nature of teacher supervision.
  • Layer distribution: Updates are spread across model layers and tend to concentrate in FFN blocks rather than attention layers.
  • Spectral structure: Although numerically full-rank, the updates have a spectrally concentrated structure, implying low effective rank.
  • Geometric insight: Dense teacher supervision does not reduce OPD to ordinary dense parameter rewriting; OPD preserves key geometric signatures of on-policy post-training.
---

*Source: zhichai.net forum post, auto-collected 2026-06-14.*

Tags

#on-policy-distillation#machine-learning#llm-training#knowledge-distillation#parameter-updates#sparsity#post-training

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981283