English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Pass the Baton: Relay-OPD Uses Teacher Takeovers to Fix On-Policy Distillation's Prefix Failure

Forum topic · ✨步子哥 · 2026-07-29

Summary

Researchers from Zhejiang University and Alibaba's Yuvion team propose Relay-OPD (Relay On-Policy Distillation), a method that fixes a structural weakness in on-policy distillation (OPD) called prefix failure: once a student model takes a wrong turn early in its reasoning chain, all subsequent tokens build on that error, and teacher feedback becomes unreliable. The key insight is an asymmetry: on failing prefixes, teachers tend to switch to reflection tokens (But, Wait, However) while students persist with continuation tokens (So, Therefore). Relay-OPD detects this divergence on the fly and briefly hands generation control to the teacher, which corrects the direction and returns control, producing relayed trajectories trained with standard OPD. Using Qwen3-4B as teacher and Qwen3-0.6B/1.7B as students across eight math benchmarks, Relay-OPD raises average accuracy to 46.96% (versus 41.23% for OPD) while cutting training trajectory length by 50.7%. The post also discusses limitations: hand-crafted reflection word sets, math-only validation, and reliance on teacher-student reflection asymmetry.

Pass the Baton: When the Student Drifts Off Course, the Teacher Briefly Takes the Wheel

Paper: Pass the Baton: Trajectory-Relayed On-Policy Distillation arXiv: 2607.26057 Authors: Haolei Xu, Xiaowen Xu, Haiwen Hong, et al. (Zhejiang University / Alibaba Yuvion) GitHub: https://github.com/zju-real/Relay-OPD

---

A Relay Race Metaphor

Imagine a 4×100 meter relay. You are a coach training a young runner for the third leg. The standard training method: let the runner complete the whole leg, then review the footage afterward — "your cadence was off here," "your arm swing was stiff there."

But there's a problem: once the runner veers off course at the start, the entire rest of the leg is run at an angle. By the time you grade them at the finish line, they may have left the track, knocked over three rows of hurdles, and tripped themselves. Your feedback is "you went the wrong way" — accurate, but useless, because the misalignment has accumulated over the whole leg.

The smarter approach? The moment they drift off course, take the baton, run a few steps to correct the heading, then hand the baton back.

This is the core idea behind Relay-OPD (Relay On-Policy Distillation) proposed by the Zhejiang University and Alibaba Yuvion team. The title's "Pass the Baton" precisely captures the mechanism.

---

The Problem: Prefix Failure

To understand why "relaying" is needed, first understand what it solves.

On-Policy Distillation (OPD) is one of the mainstream methods for compressing large-model capability into small models. The logic: let the student model generate its own reasoning trajectories, then have the teacher model assign a probability to every token of that trajectory as a supervision signal. Because the trajectories are generated by the student itself, the training and inference distributions match — this is OPD's core advantage over SFT.

But OPD has an overlooked structural flaw, which the paper calls prefix failure:

> Once the student goes wrong early in the reasoning chain, everything generated afterward is built on that deviation, producing long "misguided continuations."

A concrete example: the student solves a math problem but picks the wrong auxiliary line in step one. The next ten derivation steps may be exquisite, but they're built on a faulty foundation. The teacher gives token-level feedback on each of those ten steps — but that feedback is unreliable, because the teacher is evaluating a trajectory that should never have existed. Worse, those long trajectories burn enormous training compute for nothing.

Previous solutions all have structural defects:

  • Fixed-length truncation (ESR, FastOPD): cut at a fixed position regardless of where the reasoning actually failed. Cut too early and you waste good trajectories; cut too late and errors have already accumulated.
  • Offline rewriting (TRD): wait for the student to finish, then have the teacher rewrite it. But rewritten trajectories carry visible "editing traces," and the student never learns a naturally unfolding reasoning rhythm.
  • Token-level mixing (SKD): switch between teacher and student tokens based on distribution divergence, but the signal is too blunt — distribution divergence does not equal a wrong direction.
  • What's missing is a mechanism that can judge "online" whether the student has gone off course.

    ---

    The Key Insight: The Teacher Turns Back, the Student Doubles Down

    Relay-OPD's core finding fits in one sentence:

    > On failing prefixes, the teacher tends to turn around; the student tends to press on.

    This "teacher–student continuation asymmetry" is the foundation of the entire method. Concretely: when reasoning reaches a fork, the teacher's most likely next token is a reflection word — "But," "Wait," "However" — signaling a restart; the student's most likely next token is a continuation word — "So," "Therefore" — doubling down on the wrong direction.

    The paper's Figure 1a shows a real case: at some handoff trigger point, the teacher assigns 74.4% probability to "But," far above its probability for the student's preferred "So"; meanwhile the student gives "So" 50.6% and "But" nearly zero.

    This asymmetry can be detected without any external labels or verifiers. Just check: is the teacher's top-1 token in a predefined reflection-word set ℛ, and does the student's top-K support set contain no reflection words? If the teacher wants to reflect but the student doesn't, hand off the baton.

    More formally, the handoff criterion is:

    \[\phi(h) = \mathbb{1}[a^T(h) \in \mathcal{R}] \cdot \mathbb{1}[\mathcal{K}_S(h) \cap \mathcal{R} = \emptyset]\]

    where \(a^T(h)\) is the teacher's top-1 token at the current prefix, \(\mathcal{K}_S(h)\) is the student's top-K support set, and \(\mathcal{R}\) is the reflection-word set (Wait, But, However, etc., with case and leading-space variants).

    The elegance: it doesn't detect "the student is wrong" — it detects "the teacher would turn while the student wouldn't." These are not equivalent: the student might be wrong while the teacher is wrong too, or the student might be right while the teacher wants to reflect. Relay-OPD intervenes only on a directional disagreement, which ensures precision.

    ---

    Relay Trajectories: The Teacher Runs a Leg, the Student Carries On

    Once a handoff triggers, Relay-OPD constructs a "relay trajectory":

    1. The student runs the first segment (student leg), stopping at the handoff point. 2. The teacher takes over from that point for a short stretch (teacher leg), typically containing a reflection token plus corrective reasoning. 3. After finishing, the teacher hands control back to the student. 4. The student continues generating from the teacher's state until the trajectory ends. 5. Standard OPD training is applied to this mixed trajectory.

    The key constraint is the relay budget: only a limited number of teacher interventions per trajectory. The paper uses budget = 2 or 3. This serves two purposes:

  • Concentrating interventions at critical early positions: the earlier an error is corrected, the smaller the accumulated loss. With a limited budget, interventions naturally go to the most critical forks.
  • Limiting divergence from the student policy: if the teacher runs too much, the trajectory becomes teacher-generated, and OPD's advantage of training on the student's own distribution is lost.
  • The entire rollout runs inside a speculative decoding engine — the engineering highlight of §3.3. Teacher and student share one inference engine; teacher intervention just switches parameters, no model reloading. Relay trajectory generation costs nearly the same as single-model inference.

    ---

    The Numbers: Higher Accuracy, Shorter Trajectories

    Setup: teacher is Qwen3-4B-Instruct-2507; students are Qwen3-0.6B and 1.7B Non-Thinking versions. Eight math reasoning benchmarks: AIME24/25/26, MATH, AMC23, Olympiad Bench, HMMT Feb26, HMMT Nov25.

    Main results (Qwen3-1.7B student):

    | Method | Avg Accuracy | Training Trajectory Length | |------|-----------|-------------| | Student baseline | 24.84 | — | | SFT | 33.20 | 4262 | | KD | 33.75 | 4262 | | GRPO | 34.42 | 2558 | | OPD | 41.23 | 4658 | | TRD | 30.69 | 2785 | | FastOPD | 45.47 | 2709 | | SKD | 42.35 | 4753 | | Relay-OPD | 46.96 | 2296 |

    Key figures:

  • +5.73% over OPD, +1.49% over the strongest baseline FastOPD.
  • Training trajectory length drops from 4658 to 2296 — a 50.7% reduction. This is the direct effect of truncating prefix failures — students no longer run long erroneous continuations.
  • On AIME25 and AIME26, +7.29% and +7.19% over OPD respectively.
  • The 0.6B student shows a similar trend: +3.01% over OPD.
  • The FastOPD comparison is especially interesting. FastOPD is the strongest trajectory-intervention baseline, concentrating training signal at the sequence front via fixed-length truncation. Relay-OPD's responses on AIME25, AIME26, and HMMT Feb26 are 17.9%, 14.2%, and 28.3% shorter than FastOPD's, while accuracy is +2.39%, +4.17%, and +1.14% higher.

    Shorter and more accurate — exactly the intuition of the baton: correct early, and there's no need for a long detour later.

    ---

    Why TRD and SKD Fall Short

    The paper's comparisons are informative.

    TRD (offline rewriting) underperforms standard OPD on both students (30.69 vs 41.23 on 1.7B). Rewritten trajectories carry visible "editing traces" — the student never learns a naturally unfolding reasoning rhythm. Appendix E.2 gives a case study: TRD's rewrites read like after-the-fact patches, not natural thought.

    SKD (token-level mixing) gains only 1.12 points over OPD on 1.7B and loses 3.65 points on 0.6B. It cannot break the student's established repetitive generation patterns — Appendix E.3 shows SKD falling into repetition loops.

    Relay-OPD's advantage: it neither patches after the fact nor averages at the token level — it delivers an explicit "restart" signal at the exact moment of directional disagreement.

    ---

    Conceptual Position: Another Case of "Solving the Problem at a Different Level"

    Relay-OPD evokes a lineage previously traced in this column — "solving problems by changing levels":

  • Octopuses use RNA editing for real-time adjustments outside the DNA blueprint
  • Slime molds externalize memory in slime trails
  • Birds combine magnetite amplifiers with radical-pair sensors for magnetoreception
  • SOPHIA assigns different exit directions to different states
  • EvoThink slices reasoning into atomic units
  • Möbius RoPE changes positional encoding via topological intervention
  • Mantis shrimp use photonic crystals to selectively filter shockwaves
  • Euclid-MCP offloads reasoning to Prolog
  • ACE isolates context pollution with subagents
Relay-OPD joins this lineage: it doesn't try to make the model stronger overall — it temporarily switches control at the moment of drift. It is isomorphic to the octopus's "don't change the blueprint, change the construction plan" — leave the student's capability untouched and intervene in real time at the trajectory level.

More deeply, Relay-OPD reveals a general engineering principle: the granularity of optimization should match the granularity of the object being optimized. Reasoning is a directional trajectory process, not a pile of independent tokens. Token-level supervision (standard OPD) is too fine-grained to see direction; trajectory-level after-the-fact rewriting (TRD) is too coarse and loses naturalness. Relay-OPD intervenes at the intermediate granularity of "directional turning points" — both precise and natural.

---

Honest Assessment

Some limitations worth noting:

1. The reflection-word set ℛ is hand-crafted. Appendix A.1 lists it fully: Wait, But, However, and variants. It works well for explicitly reflective domains like math reasoning, but may not transfer to code generation or open-ended dialogue. Automatically discovering domain-specific reflection words is an open problem.

2. Validated only on math reasoning. All eight benchmarks are math problems. Prefix failure is most visible in math (long chains, severe error accumulation), but whether other tasks benefit equally remains to be seen.

3. The relay budget is a hyperparameter. Budget = 2–3 is used, with no method for adapting it to task difficulty. Easy problems may need no intervention; hard ones may need more.

4. The teacher–student asymmetry assumption. The method assumes teachers reflect on failing prefixes while students press on. This holds for the Qwen3 family, but other model families may have different reflection habits — some teachers may also double down, disabling the handoff trigger.

But as a mechanism-level insight, "the teacher turns back; the student presses on" is memorable on its own. It tells us: the difference between strong and weak models lies not just in how much they know, but in when they know to admit a mistake. The weak model's fundamental problem isn't lack of knowledge — it's not knowing when it should turn around.

---

Closing

Relay-OPD recalls an old joke: a man asks for directions, and the local answers, "If you want to get there, you shouldn't start from here."

Standard OPD is letting the student start from "here," grind to the end, and be told "you went the wrong way." Relay-OPD is the teacher taking the wheel the moment the student drifts, steering back onto the road, and handing the wheel back. What the student learns isn't "how to drive the whole route," but "which intersection to turn around at."

This insight runs deeper than Relay-OPD itself. It suggests a new philosophy of supervision: supervision shouldn't only arrive at the finish line, nor be spread evenly across every step — it should concentrate at directional turning points. Like a good driving instructor who doesn't nag the whole way, but says "change lanes here" at the critical intersection.

Paper: https://arxiv.org/abs/2607.26057 HTML version: https://arxiv.org/html/2607.26057v1 Code: https://github.com/zju-real/Relay-OPD Project page: https://zju-real.github.io/Relay-OPD

Tags

#on-policy-distillation#knowledge-distillation#llm-reasoning#relay-opd#prefix-failure#qwen3#speculative-decoding#model-training

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503776