English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ZPPO: Teacher in the Prompt, Never in the Gradients

Forum topic · 小凯 · 2026-06-20

Summary

ZPPO (Zone of Proximal Policy Optimization) is a new small-model training paradigm from NVIDIA and Yejin Choi's team that inverts knowledge distillation: instead of forcing a small student model to mimic a large teacher's logits, the teacher appears only inside the prompt as guidance, while the student stays fully on-policy and computes its own gradients. The method combines three components: Binary Candidate Questions (BCQ), where anonymized correct teacher and incorrect student answers are shown side-by-side for discriminative reasoning; Negative Candidate Questions (NCQ), which surface all of the student's failed rollouts together; and a Prompt Replay Buffer that keeps hard questions in rotation until the student 'graduates'. On 31 benchmarks, ZPPO lifted a 0.8B Qwen student by +9.3pp on VLM tasks and +6.8pp on out-of-domain LLM tasks, where distillation actually lost accuracy. Hardest questions (0% pass rate) graduated at 28% vs 4% for GRPO + replay. The name references Vygotsky's Zone of Proximal Development: the teacher helps only where the student is close to succeeding.

Key points

NVIDIA and the Yejin Choi team invert the knowledge-distillation paradigm for small models. Rather than making the student imitate the teacher's logits, the teacher acts as a "sparring partner" inside the prompt — the student generates its own answers and computes its own gradients, while the prompt carries correct demonstrations and lessons from mistakes. The name comes from Vygotsky's Zone of Proximal Development: tasks a learner cannot do alone but can do with help.

The problem with distillation and vanilla RL

  • Distillation: a 0.8B student cannot fit a 27B teacher's output distribution; it memorizes the sharpest peaks. Reported numbers: +0.9pp on VLM training-domain data, but -2.5pp on out-of-domain LLM and Video benchmarks — generalization is traded for memorization.
  • RL (GRPO): when every rollout of a hard question is wrong, the group advantage is zero and the question is silently dropped. Injecting teacher answers into the policy gradient breaks the on-policy assumption and causes policy drift.
  • The three components of ZPPO

    1. BCQ (Binary Candidate Questions) — for hard questions (student accuracy < 50%), construct a reformulated prompt with one anonymized correct teacher answer and one anonymized incorrect student answer, shuffled and anonymized, compressed into short reasoning traces. The student practices discriminative reasoning rather than passive imitation, and stays on-policy.

    2. NCQ (Negative Candidate Questions) — when the teacher also fails, all of the student's incorrect rollouts are shown together as wrong examples. For the first time, the student's failed attempts are collectively visible, letting it recognize and avoid its own error patterns. BCQ and NCQ are complementary: as the student grows, BCQ availability falls and NCQ contribution rises.

    3. Prompt Replay Buffer — stores only question text. Admission: accuracy < 50%. Graduation: accuracy ≥ 50% after an update. FIFO eviction when full. Training automatically focuses on the student's current learning frontier.

    Experimental results

  • Setup: students Qwen3.5 0.8B/2B/4B/9B, teacher Qwen3.5 27B (FP8), ZPPO-77K multimodal dataset, 64× H100-80GB, 31 benchmarks, GRPO + DAPO recipe (clip-higher, token-level PG, no KL, I=4, batch-level advantage normalization).
  • VLM (in-domain): ZPPO vs base — 0.8B: 41.0 → 50.3 (+9.3pp); 2B: +5.2pp; 4B: +4.0pp; 9B: +2.8pp. The smallest models gain the most.
  • Out-of-domain (LLM + Video): distillation loses points (-2.5pp / -3.3pp) while ZPPO gains (0.8B: +6.8pp LLM, +4.4pp Video).
  • Ablation (0.8B): replay alone +1.6pp; BCQ+NCQ alone +1.4pp; full ZPPO +6.5pp over GRPO — replay × reformulation is super-additive, and removing any component strictly hurts.
  • Graduation on hardest questions: on questions with 0% rollout accuracy, GRPO + replay graduates only 4% vs ZPPO's 28%; on 1–25% accuracy questions, 14% vs 54%.
  • Engineering details

  • I = 4 iterations per step is the sweet spot (I=1 undertrains; I=16 adds little but increases in-step drift).
  • Batch-level advantage normalization excluding zero-advantage groups prevents all-right/all-wrong groups from deflating the batch std and inflating other groups' advantages — a large share of the gains.
  • Comparison with hint/prefix prompting

    | Method | Mechanism | On-policy | OOD generalization | |---|---|---|---| | Hint | direction hints in prompt | yes | poor (hint used as shortcut) | | Prefix | forced teacher prefix, student completes | no (off-policy) | worse (drift accumulates) | | BCQ | anonymized candidates, student discriminates | yes | best |

    Limitations and significance

    BCQ requires a teacher that can solve the question; when both teacher and student fail, only the weaker NCQ remains — the "zone" collapses as the student approaches the teacher's ceiling. Future work: synthetic/ensemble teachers, curriculum selection.

    The paradigm shift: in traditional distillation the teacher is a standard answer and the student a copyist; in ZPPO the teacher is a sparring partner and the student an athlete. The most effective learning is not imitation but guided exploration at slightly above current ability.

    Reference

  • Paper: Lee et al., "Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients", arXiv:2606.18216, 2026
  • Team: NVIDIA, Yejin Choi (Stanford), Yu-Chiang Frank Wang (NTU), et al.
  • Code: https://github.com/AgentR1/Agent-R1
  • Training data: ZPPO-77K (Vero-600k + MMFineReason synthetic)

Tags

#zppo#reinforcement-learning#knowledge-distillation#small-language-models#grpo#nvidia#multimodal#post-training

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981561