Tagged articles

mode seeking

1 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 9, 2026 · Artificial Intelligence

Why Online RL (OPD, RLHF) Uses Reverse KL While Offline SFT Distillation Prefers Forward KL

The article explains that forward KL encourages a student model to cover all major teacher modes, whereas reverse KL seeks a single dominant mode, and shows why online reinforcement learning methods like OPD and RLHF adopt reverse KL while offline SFT distillation relies on forward KL.

KL DivergenceOPDRLHF
0 likes · 7 min read
Why Online RL (OPD, RLHF) Uses Reverse KL While Offline SFT Distillation Prefers Forward KL