Machine Learning Algorithms & Natural Language Processing
Aug 24, 2026 · Artificial Intelligence
When Online Distillation Goes Off‑Track: How Relay‑OPD Lets the Teacher Take Over at Critical Moments
The article analyzes the prefix‑failure problem in on‑policy distillation, introduces Relay‑OPD with a handoff trigger that lets a teacher model intervene locally, and shows through eight math‑reasoning benchmarks that this approach improves accuracy by up to 7.3% while cutting training trajectory length by more than half.
Relay-OPDlarge language modelsmath reasoning benchmarks
0 likes · 13 min read
