Why Online RL (OPD, RLHF) Uses Reverse KL While Offline SFT Distillation Prefers Forward KL
The article explains that forward KL encourages a student model to cover all major teacher modes, whereas reverse KL seeks a single dominant mode, and shows why online reinforcement learning methods like OPD and RLHF adopt reverse KL while offline SFT distillation relies on forward KL.
