Relay-OPD: Teacher Steps In at Critical Moments to Fix Student Reasoning Errors in Online Distillation
Zhejiang University and Alibaba propose Relay-OPD, an online distillation method that detects student reasoning errors via teacher-student token divergence and briefly hands control to the teacher for targeted correction, achieving +5.73% accuracy on math benchmarks while halving training trajectory length.
As large language models grow stronger, efficiently transferring their capabilities to smaller models during post-training becomes crucial. On-policy distillation (OPD) lets a student model learn from a teacher on the student's own generated trajectories, reducing distribution shift between training and inference. However, OPD suffers from prefix failure : once the student makes an early mistake, all subsequent generation builds on that error, creating long, unreliable trajectories that waste compute.
Limitations of Existing Approaches
Fixed-length truncation (ESR, FastOPD) cuts rollouts at predetermined positions, unrelated to where reasoning actually fails.
Offline rewriting (TRD) waits until the full trajectory is generated before teacher correction; intervention is too late and leaves visible rewrite artifacts.
Token-level mixing (SKD) switches based on general teacher-student distribution divergence, lacking an explicit signal that the reasoning direction has failed.
In short, current methods either cut inaccurately, correct too late, or switch blindly. A mechanism that triggers online and is driven by the reasoning state itself is missing.
Key Observation: Teacher-Student Divergence at Failure Points
The authors notice that on a deviated prefix, the teacher and student exhibit distinct continuation instincts. The teacher tends to pause, reflect, and pivot — its top next token is often a reflection word like Wait, But, or However. The student, however, tends to continue along the wrong path (e.g., So). In a concrete example, at a deviated position the teacher assigned 74.4% probability to But while the student assigned 50.6% to So. This divergence is directly observable during generation without any verifier, process labels, or reward model. The paper calls such positions handoff triggers .
Trajectory Intervention Experiments
Two critical findings emerge from controlled interventions:
Correction can be extremely local : replacing only the single reflection token at each trigger (teacher tokens constitute just 0.35% of all tokens) lifts accuracy from 27.73 to 34.96 (+7.23%).
Intervention value is highly front-loaded : keeping intervention length fixed but moving the intervention point later causes accuracy to drop from 41.99 to 33.98 to 29.49. As the prefix grows, the teacher is also led astray by the student's context, narrowing the teacher-student gap; late takeovers become ineffective.
Moreover, extending teacher takeover length yields diminishing returns: increasing teacher token share from 17.52% to 28.52% only moves accuracy within 41–44, far below the teacher's standalone 60.55. Intervention must be timely and restrained.
Relay-OPD: Three Core Designs
1. Label-Free Handoff Trigger
Predefine a reflection vocabulary R = {Wait, But, However, and case/space variants}. At each student generation step, if the teacher's top-1 token belongs to R while the student's top- K (main experiment K =5) contains no reflection word, a handoff is triggered.
2. Relay Budget (M, L)
Each trajectory allows at most M teacher takeovers. Each takeover starts at the triggered reflection token and continues for L natural paragraphs (delimited by \n\n, averaging ~23.2 tokens). Measuring by paragraphs ensures each takeover ends at a structurally complete reasoning unit. After a teacher segment, if budget remains, the student resumes from the corrected prefix; after the M -th takeover the rollout terminates. Main experiment uses ( M , L ) = (2, 3). Unlike fixed truncation, both intervention and termination positions are state-dependent.
3. Training Objective for Relay Trajectories
The authors argue that teacher guidance on relay trajectories is not equivalent to the teacher's ideal behavior on its own trajectories — earlier experiments showed even with ~30% teacher tokens, accuracy remains far below teacher standalone. Therefore, the student should not blindly mimic the teacher's full output distribution. Relay-OPD performs single-sample distillation directly on the actually generated relay tokens, selectively absorbing correction signals rather than fitting the teacher's complete distribution.
Engineering: Unified Speculative Decoding Engine
Relay-OPD requires the teacher's next-token distribution at every student generation step to evaluate the trigger, and the two models must alternate on the same trajectory. Maintaining two separate inference engines would incur heavy synchronization and context-switching overhead. The solution: fold the entire relay process into a single speculative decoding engine, with the student as the draft model and the teacher as the target model. During student segments, the target is the student itself (all drafts accepted, equivalent to normal student decoding). During teacher segments, standard speculative rejection sampling applies. The teacher logits computed for verification simultaneously provide the trigger signal at zero extra cost, and the speculative sampling correctness guarantee ensures the single-engine trajectory distribution matches true alternating generation.
Experimental Results
Eight Math Benchmarks: Comprehensive Gains, Trajectory Length Halved
Teacher: Qwen3-4B-Instruct-2507; Students: Qwen3-0.6B and Qwen3-1.7B (Non-Thinking). Training on DAPO-Math-17K English subset using verl and vLLM. No verifier, process supervision labels, or answer correctness labels used. Evaluated on AIME 2024/2025/2026, MATH500, AMC 2023, OlympiadBench, HMMT Feb 2026 and Nov 2025 (eight benchmarks).
Relay-OPD achieves best or second-best on all eight benchmarks for both student sizes.
1.7B student: average accuracy 46.96 vs. standard OPD 41.23 (+5.73%). AIME 2025 +7.29%, AIME 2026 +7.19%. Outperforms strongest trajectory-intervention baseline FastOPD (45.47) by +1.49%.
0.6B student: +3.01% over OPD, slightly beats FastOPD (+0.62%).
Baseline failure modes: TRD underperforms OPD due to unnatural rewrite traces; SKD cannot break student's repetitive generation patterns; FastOPD concentrates signal early but cannot demonstrate recovery from a failed prefix.
Efficiency
1.7B student average training trajectory length: 2,296 tokens vs. OPD 4,658 (50.7% reduction) and FastOPD 2,709.
Convergence in 35 steps vs. OPD 55 steps.
0.6B student trajectory length reduced 63.9% vs. OPD.
Shorter Inference, Higher Accuracy
Compared to FastOPD, Relay-OPD reduces average response length on AIME 2025, AIME 2026, HMMT Feb 2026 by 17.9%, 14.2%, 28.3% respectively, while improving accuracy by +2.39%, +4.17%, +1.14%. Pass@k under various sampling budgets also uniformly exceeds standard OPD, indicating the student learns a reasoning habit of pivoting before errors snowball.
Adaptive Decay of Teacher Intervention
Training dynamics show the proportion of trajectories using the full relay budget drops from 75–85% initially to 50–60%. Teacher token share falls from ~13% to a stable 2–3% after ~20 steps. As the student improves, fewer handoffs are needed, and intervention intensity converges automatically without manual scheduling. Policy entropy remains higher than OPD and FastOPD, reflecting increased exploration from teacher-induced prefix changes.
Ablation: Teacher Segments and Training Objective Are Both Essential
Teacher segments necessary? With M =1, "trigger-then-stop" (detect trigger and truncate, no teacher tokens) yields 43.48; adding L =3 teacher segment raises to 46.25 (+2.77%). Gain comes from corrected context and local reasoning demonstration, not just dynamic truncation.
Training objective choice? Distilling on actual relay tokens (46.96) outperforms using student draft tokens (44.56) and forward KL to teacher's full distribution (44.08). The latter forces the student to absorb all teacher outputs, including unreliable guidance.
Sensitivity Analysis
L from 0 to 4: accuracy rises from 44.31 to 47.10; L =5 starts to decline.
M =2 optimal (46.96); M =4 drops to 44.01.
Overly long or frequent teacher takeovers push trajectories too far from the student's own policy, undermining the on-policy advantage. Moderate intervention is the method's essence.
The trigger signal comes from direct teacher-student logit comparison, implemented as a byproduct of the speculative decoding engine, requiring no verifier or annotations. Subsequent work in agent distillation has already adopted and extended this handoff mechanism (arXiv:2608.01953). Training, ablation, and evaluation code are open-sourced at https://github.com/ZJU-REAL/Relay-OPD.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
