Key points
NVIDIA and the Yejin Choi team invert the knowledge-distillation paradigm for small models. Rather than making the student imitate the teacher's logits, the teacher acts as a "sparring partner" inside the prompt — the student generates its own answers and computes its own gradients, while the prompt carries correct demonstrations and lessons from mistakes. The name comes from Vygotsky's Zone of Proximal Development: tasks a learner cannot do alone but can do with help.
The problem with distillation and vanilla RL
- Distillation: a 0.8B student cannot fit a 27B teacher's output distribution; it memorizes the sharpest peaks. Reported numbers: +0.9pp on VLM training-domain data, but -2.5pp on out-of-domain LLM and Video benchmarks — generalization is traded for memorization.
- RL (GRPO): when every rollout of a hard question is wrong, the group advantage is zero and the question is silently dropped. Injecting teacher answers into the policy gradient breaks the on-policy assumption and causes policy drift.
- Setup: students Qwen3.5 0.8B/2B/4B/9B, teacher Qwen3.5 27B (FP8), ZPPO-77K multimodal dataset, 64× H100-80GB, 31 benchmarks, GRPO + DAPO recipe (clip-higher, token-level PG, no KL, I=4, batch-level advantage normalization).
- VLM (in-domain): ZPPO vs base — 0.8B: 41.0 → 50.3 (+9.3pp); 2B: +5.2pp; 4B: +4.0pp; 9B: +2.8pp. The smallest models gain the most.
- Out-of-domain (LLM + Video): distillation loses points (-2.5pp / -3.3pp) while ZPPO gains (0.8B: +6.8pp LLM, +4.4pp Video).
- Ablation (0.8B): replay alone +1.6pp; BCQ+NCQ alone +1.4pp; full ZPPO +6.5pp over GRPO — replay × reformulation is super-additive, and removing any component strictly hurts.
- Graduation on hardest questions: on questions with 0% rollout accuracy, GRPO + replay graduates only 4% vs ZPPO's 28%; on 1–25% accuracy questions, 14% vs 54%.
- I = 4 iterations per step is the sweet spot (I=1 undertrains; I=16 adds little but increases in-step drift).
- Batch-level advantage normalization excluding zero-advantage groups prevents all-right/all-wrong groups from deflating the batch std and inflating other groups' advantages — a large share of the gains.
- Paper: Lee et al., "Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients", arXiv:2606.18216, 2026
- Team: NVIDIA, Yejin Choi (Stanford), Yu-Chiang Frank Wang (NTU), et al.
- Code: https://github.com/AgentR1/Agent-R1
- Training data: ZPPO-77K (Vero-600k + MMFineReason synthetic)
The three components of ZPPO
1. BCQ (Binary Candidate Questions) — for hard questions (student accuracy < 50%), construct a reformulated prompt with one anonymized correct teacher answer and one anonymized incorrect student answer, shuffled and anonymized, compressed into short reasoning traces. The student practices discriminative reasoning rather than passive imitation, and stays on-policy.
2. NCQ (Negative Candidate Questions) — when the teacher also fails, all of the student's incorrect rollouts are shown together as wrong examples. For the first time, the student's failed attempts are collectively visible, letting it recognize and avoid its own error patterns. BCQ and NCQ are complementary: as the student grows, BCQ availability falls and NCQ contribution rises.
3. Prompt Replay Buffer — stores only question text. Admission: accuracy < 50%. Graduation: accuracy ≥ 50% after an update. FIFO eviction when full. Training automatically focuses on the student's current learning frontier.
Experimental results
Engineering details
Comparison with hint/prefix prompting
| Method | Mechanism | On-policy | OOD generalization | |---|---|---|---| | Hint | direction hints in prompt | yes | poor (hint used as shortcut) | | Prefix | forced teacher prefix, student completes | no (off-policy) | worse (drift accumulates) | | BCQ | anonymized candidates, student discriminates | yes | best |
Limitations and significance
BCQ requires a teacher that can solve the question; when both teacher and student fail, only the weaker NCQ remains — the "zone" collapses as the student approaches the teacher's ceiling. Future work: synthetic/ensemble teachers, curriculum selection.
The paradigm shift: in traditional distillation the teacher is a standard answer and the student a copyist; in ZPPO the teacher is a sparring partner and the student an athlete. The most effective learning is not imitation but guided exploration at slightly above current ability.