OPD Evolution: From CoT SFT to Self‑Distillation and Preference Optimization

Since 2026, On‑Policy Distillation (OPD) has rapidly become a focal research area, evolving from offline teacher‑generated data to online student‑driven supervision, with advances such as OPD+, Direct OPD, weak‑to‑strong OPD, self‑distillation techniques, and preference‑optimization signals reshaping post‑training for large language models.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
OPD Evolution: From CoT SFT to Self‑Distillation and Preference Optimization

Starting in 2026, works labeled On‑Policy Distillation (OPD) have proliferated in the post‑training of inference models, marking the next hot topic after Long CoT distillation. Unlike traditional offline distillation, OPD lets the student receive teacher supervision on trajectories sampled by its own policy, directly addressing the teacher‑student mismatch (ESR rethinking OPD).

OPD+ redesigns the core advantage term in the objective function. The previous teacher‑probability‑ratio advantage showed efficiency bottlenecks for capability transfer; by explicitly constructing a reward component, OPD+ markedly improves the transfer efficiency from teacher to base model.

Direct OPD challenges the implicit assumption that the teacher must be stronger than the student. It enables a weak model to distill a strong model on its own rollouts, eliminating the computational cost of regenerating data after each model upgrade and demonstrating the feasibility of OPD as a continual improvement mechanism in weak‑to‑strong generalization scenarios.

The subsequent weak‑to‑strong OPD series further expands the teacher pool to weak or homogeneous teachers, systematically asking which teacher, in what manner, and on which trajectories yields the most effective student improvement.

Viewed historically, OPD is the inevitable result of the long‑term evolution of distillation techniques. The offline stage, exemplified by DeepSeek‑R1’s data distillation and ReasonLite’s two‑stage long/short CoT distillation, proved that data quality outweighs quantity. The online stage, represented by ESR, OPD+, and the weak‑to‑strong series, shifts training data from “what the teacher can provide” to “what the student needs,” opening new space in inference‑length control, transfer efficiency, and continuous self‑improvement.

Three research lines converge within this evolution: (1) Long/short CoT supervised fine‑tuning (SFT), where 1,000 carefully selected examples can trigger chain‑of‑thought reasoning and techniques like TokenSqueeze and ConPress show that long chains can be compressed without accuracy loss; (2) Self‑distillation/self‑evolution, with SPIN introducing self‑play fine‑tuning, R‑Zero and Absolute Zero demonstrating zero‑data self‑evolution, yet still facing reward‑hacking and instability challenges (Can LRMs Self‑Train); (3) DPO / preference optimization, answering “what signal to train with” rather than “what data to train with,” progressing from offline variants such as SimPO and ORPO to online signals like OAIF and Step‑DPO, which merge directly with OPD.

Collectively, these lines indicate a clear trend: post‑training of reasoning models is moving from “offline, static, teacher‑driven” to “online, dynamic, student‑driven.” OPD serves as the theoretical anchor of this shift, unifying data source, supervision signal, and training paradigm, and clarifying the positions of long/short CoT SFT, self‑distillation, and preference optimization within the overall taxonomy.

The report surveys roughly 120 verified papers from January 2025 to August 2026, maps the three sub‑directions within the OPD lineage, highlights remaining research gaps, and offers concrete topic suggestions for future work.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelsNLPpreference optimizationself-distillationOn-Policy DistillationOPD
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.