Tagged articles

Advantage Estimation

2 articles · Page 1 of 1
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Sep 29, 2026 · Interview Experience

GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking

This article breaks down six high-frequency GRPO interview questions from top Chinese tech companies, covering GRPO vs PPO trade-offs, group-relative advantage calculation with concrete numbers, handling all-correct/all-wrong sample groups, KL constraint mechanics, convergence monitoring priorities, and reward hacking detection via shadow evaluation.

Advantage EstimationGRPOInterview Preparation
0 likes · 19 min read
GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking
Baobao Algorithm Notes
Baobao Algorithm Notes
Nov 18, 2024 · Artificial Intelligence

Demystifying Actor‑Critic and PPO: From Policy Gradients to Practical RL

This article provides a thorough, step‑by‑step explanation of reinforcement‑learning theory—covering policy‑based objectives, value‑function definitions, the derivation of policy gradients, actor‑critic architecture, advantage estimation, importance sampling, GAE, and the PPO algorithm—aimed at readers with little prior RL knowledge.

Advantage EstimationPPOactor-critic
0 likes · 31 min read
Demystifying Actor‑Critic and PPO: From Policy Gradients to Practical RL