Tagged articles

KL Divergence

17 articles · Page 1 of 1
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Sep 29, 2026 · Interview Experience

GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking

This article breaks down six high-frequency GRPO interview questions from top Chinese tech companies, covering GRPO vs PPO trade-offs, group-relative advantage calculation with concrete numbers, handling all-correct/all-wrong sample groups, KL constraint mechanics, convergence monitoring priorities, and reward hacking detection via shadow evaluation.

Advantage EstimationGRPOInterview Preparation
0 likes · 19 min read
GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking
Geek Labs
Geek Labs
Sep 9, 2026 · Artificial Intelligence

Heretic: Automated Model Surgery Removes Refusal Without Damaging Capabilities

This article explains how Heretic automates directional ablation to permanently remove refusal behavior from language models by editing model weights, contrasting it with jailbreaks and steering vectors, and details its use of Optuna optimization with KL divergence to preserve model capabilities, situating it within a broader taxonomy of model modification techniques.

AbliterationHereticKL Divergence
0 likes · 18 min read
Heretic: Automated Model Surgery Removes Refusal Without Damaging Capabilities
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 9, 2026 · Artificial Intelligence

Why Online RL (OPD, RLHF) Uses Reverse KL While Offline SFT Distillation Prefers Forward KL

The article explains that forward KL encourages a student model to cover all major teacher modes, whereas reverse KL seeks a single dominant mode, and shows why online reinforcement learning methods like OPD and RLHF adopt reverse KL while offline SFT distillation relies on forward KL.

KL DivergenceOPDRLHF
0 likes · 7 min read
Why Online RL (OPD, RLHF) Uses Reverse KL While Offline SFT Distillation Prefers Forward KL
DeepHub IMBA
DeepHub IMBA
Aug 8, 2026 · Artificial Intelligence

Why Parallel Loop Transformers Peak at Two Iterations – Insights from LoopCoder‑v2

The LoopCoder‑v2 study shows that Parallel Loop Transformers achieve their best code‑generation performance with two refinement loops, as additional loops increase memory cost without improving results and even cause performance degradation, a finding explained through detailed metric analysis and cost‑benefit reasoning.

G-SWAKL DivergenceLoopCoder-v2
0 likes · 14 min read
Why Parallel Loop Transformers Peak at Two Iterations – Insights from LoopCoder‑v2
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 16, 2026 · Artificial Intelligence

SFT, DAgger, Offline RL, and OPD: Four Methods Mapped onto a Single 2×2 Grid

The paper shows that SFT, DAgger, offline RL and OPD are the four orthogonal combinations of prefix source (teacher vs. student) and KL direction (forward vs. reverse), exposing three hidden trade‑offs—KL direction, prefix source, and training length—and proposes KL‑mixing and entropy‑gated length curricula that boost Avg@k by 3.6 points, raise Pass@k by up to 5.8 points, and cut response length by three‑fold.

DAggerKL DivergenceLLM distillation
0 likes · 17 min read
SFT, DAgger, Offline RL, and OPD: Four Methods Mapped onto a Single 2×2 Grid
Machine Heart
Machine Heart
May 29, 2026 · Artificial Intelligence

DiffusionOPD: A New Online Policy Distillation Paradigm for Multi‑Task Diffusion Models

DiffusionOPD introduces a unified on‑policy distillation framework for diffusion models that decouples single‑task online policy exploration from multi‑task capability integration, training expert teachers per task and distilling their skills into a single student model, achieving faster convergence and higher performance across composition, OCR, and aesthetic tasks.

KL DivergenceOn-Policy DistillationPPO
0 likes · 8 min read
DiffusionOPD: A New Online Policy Distillation Paradigm for Multi‑Task Diffusion Models
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Apr 12, 2026 · Artificial Intelligence

Deep Dive into Forward vs Reverse KL Divergence: When to Use Which?

The article explains the definitions, properties, and asymmetric nature of KL divergence, compares Forward KL (mean‑seeking) and Reverse KL (mode‑seeking) through bimodal examples, and provides practical guidelines for choosing between them based on sampling and probability‑evaluation capabilities in machine‑learning tasks.

KL Divergenceforward KLmachine learning
0 likes · 10 min read
Deep Dive into Forward vs Reverse KL Divergence: When to Use Which?
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Feb 22, 2026 · Artificial Intelligence

What Is On-Policy Distillation? A Deep Dive into On-Policy and Self-Distillation

The article explains On-Policy Distillation, derives its forward and reverse KL gradients, introduces Self‑Distillation where the policy serves as its own teacher, discusses practical implementation tricks such as extra‑knowledge injection, EMA or trust‑region teacher stabilization, and highlights benefits like reduced catastrophic forgetting, fewer Aha moments, and a narrower train‑test gap, especially for larger models.

EMAKL DivergenceOn-Policy Distillation
0 likes · 6 min read
What Is On-Policy Distillation? A Deep Dive into On-Policy and Self-Distillation
AI Algorithm Path
AI Algorithm Path
May 10, 2025 · Artificial Intelligence

Master KL Divergence: Definitions, Properties, and Real‑World Applications

This article explains the Kullback‑Leibler (KL) divergence for discrete and continuous distributions, outlines its non‑negativity and asymmetry, walks through a uniform‑distribution example, provides a simple Python demonstration, and discusses key applications in variational autoencoders, reinforcement‑learning policy optimization, and other machine‑learning contexts.

KL DivergenceVariational AutoEncoderinformation theory
0 likes · 7 min read
Master KL Divergence: Definitions, Properties, and Real‑World Applications
Code DAO
Code DAO
May 6, 2022 · Fundamentals

Information Theory Foundations for Machine Learning and Deep Learning

The article explains Shannon information content, entropy, cross‑entropy, KL‑divergence, conditional entropy and mutual information, illustrating each concept with coin‑flip and dice examples, visual formulas, and discusses their roles as loss functions and evaluation metrics in machine‑learning models.

Cross-EntropyKL Divergenceentropy
0 likes · 8 min read
Information Theory Foundations for Machine Learning and Deep Learning
Code DAO
Code DAO
Dec 20, 2021 · Artificial Intelligence

Exploring Latent Space with a Variational Autoencoder in TensorFlow

This article explains the theory behind variational autoencoders, details their KL‑divergence loss, provides a complete TensorFlow implementation, and demonstrates reconstruction, latent‑space visualization, and novel image generation through sampling and interpolation.

KL DivergencePythonTensorFlow
0 likes · 13 min read
Exploring Latent Space with a Variational Autoencoder in TensorFlow
Code DAO
Code DAO
Dec 10, 2021 · Artificial Intelligence

Understanding Variational Autoencoders: From Dimensionality Reduction to Generative Modeling

This article explains the principles of variational autoencoders, starting with dimensionality reduction techniques such as PCA and standard autoencoders, highlighting their limitations for data generation, and then detailing VAE's regularized latent space, variational inference, re‑parameterization, and loss formulation.

Generative ModelsKL DivergenceVAE
0 likes · 18 min read
Understanding Variational Autoencoders: From Dimensionality Reduction to Generative Modeling
21CTO
21CTO
Feb 7, 2018 · Artificial Intelligence

Demystifying Entropy: From Basic Concepts to Cross‑Entropy and KL Divergence

This article explains entropy, joint entropy, conditional entropy, and related measures such as KL divergence and cross‑entropy, using intuitive coin‑flip examples and mathematical formulas to show how they quantify uncertainty and information in probability distributions.

Cross-EntropyKL Divergenceentropy
0 likes · 14 min read
Demystifying Entropy: From Basic Concepts to Cross‑Entropy and KL Divergence
Qunar Tech Salon
Qunar Tech Salon
Mar 14, 2015 · Artificial Intelligence

Common Distance and Similarity Measures in Machine Learning and Data Mining

This article reviews the most frequently used distance and similarity formulas in machine learning and data mining, explaining their definitions, mathematical properties, practical examples, and when each metric is appropriate for measuring differences between data points or probability distributions.

KL DivergenceMahalanobis distancecosine similarity
0 likes · 13 min read
Common Distance and Similarity Measures in Machine Learning and Data Mining