Tagged articles

LLM Post-Training

4 articles · Page 1 of 1
Machine Heart
Machine Heart
Oct 2, 2026 · Artificial Intelligence

AIBuildAI PostTrain Agent: Autonomous LLM Post-Training Agent Tops PostTrainBench

AIBuildAI's open-source PostTrain Agent autonomously designs and executes LLM post-training pipelines—including data curation, algorithm selection, and hyperparameter tuning—achieving 46.6 on PostTrainBench, surpassing all frontier models and agents, nearing human expert performance (51.1) within 10 hours on a single H100.

AIBuildAIAutonomous AI ResearchKnowledge System
0 likes · 27 min read
AIBuildAI PostTrain Agent: Autonomous LLM Post-Training Agent Tops PostTrainBench
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Sep 29, 2026 · Interview Experience

GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking

This article breaks down six high-frequency GRPO interview questions from top Chinese tech companies, covering GRPO vs PPO trade-offs, group-relative advantage calculation with concrete numbers, handling all-correct/all-wrong sample groups, KL constraint mechanics, convergence monitoring priorities, and reward hacking detection via shadow evaluation.

Advantage EstimationGRPOInterview Preparation
0 likes · 19 min read
GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking
PaperAgent
PaperAgent
Sep 17, 2026 · Artificial Intelligence

LLM Post-Training Paradigm Shift: 5 New Paths Replacing Monolithic RL

This article analyzes five fundamental shifts in LLM post-training over the past six months: moving from monolithic RL to expert distillation (MOPD), online distillation as a 10x cheaper RL alternative, refined RLVR techniques addressing entropy collapse and exploration, SFT-RL distribution alignment, and data quality as an irrecoverable hard constraint.

Data QualityLLM Post-TrainingMiMo-V2-Flash
0 likes · 9 min read
LLM Post-Training Paradigm Shift: 5 New Paths Replacing Monolithic RL
Data Party THU
Data Party THU
Apr 20, 2026 · Artificial Intelligence

Can AI Rewrite Its Own Evolution Engine? Inside HyperAgents' Self‑Modification Breakthrough

The article analyzes the HyperAgents framework (DGM‑H), showing how merging task and meta agents enables metacognitive self‑modification, improves performance across coding and non‑coding benchmarks, automatically builds supporting infrastructure, and raises new safety and industry‑impact considerations.

AI safetyHyperagentsLLM Post-Training
0 likes · 11 min read
Can AI Rewrite Its Own Evolution Engine? Inside HyperAgents' Self‑Modification Breakthrough