AIBuildAI PostTrain Agent: Autonomous LLM Post-Training Agent Tops PostTrainBench
AIBuildAI's open-source PostTrain Agent autonomously designs and executes LLM post-training pipelines—including data curation, algorithm selection, and hyperparameter tuning—achieving 46.6 on PostTrainBench, surpassing all frontier models and agents, nearing human expert performance (51.1) within 10 hours on a single H100.
Overview
AIBuildAI has released PostTrain Agent , an open-source recursive self-improvement (RSI) agent that autonomously performs the entire post-training process for large language models. Given a base model, a target capability, and a compute budget, the agent designs post-training algorithms, collects and curates data, writes training code, runs experiments, and iteratively improves the model without human intervention. On the PostTrainBench benchmark (Rank et al., ICML 2026), the agent scores 46.6 , ranking first among all frontier models and agent systems, and approaching the human expert baseline of 51.1.
Why Automated Post-Training?
Post-training transforms a base model into one that follows instructions, reasons, uses tools, and understands specific domains. It involves coupled decisions: data selection and mixing, choice of supervised fine-tuning (SFT), preference optimization, or reinforcement learning, and hyperparameters such as learning rate, training steps, and checkpoint selection. These decisions depend on each other—suitable algorithms depend on data mix, which in turn depends on the base model. Traditionally, an experienced team spends weeks and significant compute to deliver a post-trained model. AIBuildAI's research question: can AI fully automate this end-to-end?
PostTrainBench: The Benchmark
PostTrainBench formalizes the task: given a base model (e.g., Qwen3-4B-Base) and a target capability (e.g., tool use), the agent runs on a single H100 for 10 hours, autonomously designing and executing a post-training recipe. The resulting checkpoint is evaluated on held-out test sets across 4 base models and 7 task categories: math reasoning (AIME 2025, GSM8K), writing (ArenaHard Writing), general QA (GPQA), health consulting (HealthBench), code (HumanEval), and function calling (BFCL). All choices—external data, mixing strategies, SFT vs. preference optimization vs. RL, hyperparameters, and training frameworks—are made by the agent.
From Linear Search to Experiment Trees
Baseline approaches use a coding agent (e.g., Claude Code) in a single CLI session: linear search where each attempt modifies the previous one, committing to a single direction early and unable to explore diverse ideas in parallel. AIBuildAI's prior systems (AIBuildAI Science, AIBuildAI-2.5) introduced whole-experiment tree search : each run is a tree of experiments; a Designer agent seeds diverse starting strategies; an Experimenter agent runs each experiment end-to-end and proposes modifications; a Selector ranks candidates by expected gain, evidence, novelty, and cost.
However, directly applying tree search to post-training revealed two limitations:
Lack of domain knowledge : Experimenters relied on parametric knowledge, leading to errors such as selecting methods without library implementations, using datasets overlapping with eval benchmarks, or applying hyperparameters from different model scales—each error costing a full training run to discover.
Fixed search topology before understanding the task : Tree nodes must be scored models, but stages like data generation produce no model or score; merging multiple trained models (fan-in) cannot be expressed.
Knowledge System: Evidence-Based Decisions
The knowledge system is a remote retrieval knowledge base served via MCP, containing four sub-libraries of knowledge files :
Workflows : 13 ordered steps of a post-training run, 34 files, each specifying judgments to settle, a concrete example, and applicability boundaries.
Methods : 111 files covering SFT, DPO, GRPO, distillation, etc., with requirements, cost, parameters, defaults, and library implementations.
Datasets : 215 files with licenses, splits, loading code, sample rows, and flags for eval benchmarks requiring isolation.
Frameworks : 108 files on when to choose a framework, how to launch, monitor, and save a training run.
Each file is sourced from public papers, library code, dataset cards, and public training logs, manually verified for scope, license, and splits, with every claim backed by a URL and date. Files mentioning eval benchmarks are excluded from the corpus. During a run, the Experimenter queries the knowledge base with the task, current plan, and issues, retrieves matching files, decides next steps, and writes observations back after each experiment, growing the knowledge base.
These files encode actionable judgments, not tutorials. For example, workflow step 1 requires the agent to understand what the scoring function rewards (final answer vs. full response, format penalties, reasoning credit) and to run the official evaluator on the untrained base model to establish a baseline and measure per-evaluation cost. Method files provide applicability conditions, objective functions, worked numerical examples, required data shapes, auxiliary models, and library implementations.
Meta-Search: Search Topology as Agent Output
Different tasks need different search shapes. Linear search suits fixed recipes; tree search suits choosing among a few plausible strategies. But post-training often involves structures neither can express. A typical recipe: prepare two training datasets (task-provided and externally built), run SFT on each on the base model, pick the better on a validation set as current best ; then enter an RL loop: from current best, run one GRPO round, evaluate on validation, replace current best only if score improves, repeat until budget exhausted.
This cannot fit in a tree because data construction yields no model or score, and fan-in (merging models) is inexpressible. AIBuildAI's solution is meta-search : a Meta Agent first researches the task, reads design knowledge files (37 runnable search patterns: chains, fan-out/fan-in, loops, phased pipelines, checkpoint tournaments, SFT-then-GRPO, budget-constrained multi-round training), queries the knowledge base, and writes a search program . The program's steps can be LLM agents, deterministic programs, or nested sub-searches, forming any execution graph—phases, tournaments, fan-in, loops. The example above becomes a few lines of code: two SFT branches, pick best, loop GRPO while budget allows.
The search program composes three primitives: Agent (LLM session with tools and typed output), Program (deterministic computation under resource limits), and Search (orchestrates Agents, Programs, and nested Searches). The 37 patterns are a vocabulary, not a whitelist; the Meta Agent can combine, modify, or invent new structures. The program need not be fixed upfront: the Meta Agent can write only the current phase (e.g., corpus construction or one training+evaluation round), execute it, read remaining budget, and hand off to a next-generation Meta Agent that writes the next phase based on empirical results.
Experimental Analysis
All experiments follow PostTrainBench rules: each task runs fully automatically on a single H100 for 10 hours; all agent roles run on Claude Opus 5 ; all results come from meta-search. Comparison numbers are from the public PostTrainBench leaderboard.
Overall Score: Beating All Frontier Models and Agents
AIBuildAI PostTrain Agent achieves a weighted aggregate score of 46.6 , higher than all frontier models and agent entries, second only to human experts (51.1). The next best are Locus (45.6) and Claude Code with Claude Fable 5 (41.8), followed by GPT-5.6 (36.2), Claude Opus 5 (35.0), Claude Opus 4.8 (33.8), Kimi K3 (32.0), GLM-5.2 (31.7), Gemini 3.1 Pro (22.0). Untrained base models average 7.5—meaning the agent adds ~39 points in 10 hours.
Same Base Model, Different System: +11.6 Points
A more telling comparison: AIBuildAI's agents all run on Claude Opus 5, while Claude Code on the same Opus 5 scores only 35.0—a gap of 11.6 points attributable to the system, not the model. AIBuildAI leads on six of seven tasks, ties on GPQA. Largest gains: BFCL (+94.3), ArenaHard Writing (+12.1), HumanEval (+9.2). Notably, AIBuildAI (Opus 5) at 46.6 also beats Claude Code running the stronger Claude Fable 5 (41.8): the knowledge system and meta-search provide more gain than upgrading the base model one generation.
Per-Task Analysis: Three Tasks Achieve Highest Agent Scores
Against all agents on the leaderboard, AIBuildAI tops AIME 2025 (15.8 vs. 13.3 for Claude Fable 5 and 9.4 for Locus), BFCL (95.8), and HumanEval (69.0 vs. Locus 66.6). On the other four tasks, gaps to the leaders are under 2 points.
BFCL : Scores depend on exact function-call parsing and schema compliance; format errors yield zero. Claude Code on Opus 5 averages 1.5, same as untrained base. AIBuildAI uses the knowledge system to pick the right training framework and format-matched datasets, scoring ≥94 on all four base models, average 95.8—the only task exceeding the human reference (85.0).
ArenaHard Writing : Pairwise preference judging reduces to data selection. AIBuildAI's meta-search designs effective preference optimization pipelines.
AIME : Answers are verifiable; sample-verify-retrain loops suit meta-search. AIME carries the highest weight in aggregate and is where all agents lag humans most. AIBuildAI achieves the highest or tied-highest agent score on every base model, including the toughest gemma-3-4b-pt (3.3 vs. ≤1.1 for others).
HumanEval : AIBuildAI scores highest on all four base models.
HealthBench : Uses physician-written rubrics; requires preference optimization that avoids rewarding verbosity.
Two Case Studies: Step-by-Step Score Improvement
In both cases, the agent deploys an open-weight teacher model on its own GPU to generate training data, and plans one phase at a time based on measured results.
HealthBench × Qwen3-4B-Base (base 13.4 → 48.6, Claude Code 40.9): Teacher generates ~48k replies, SFT yields 46.2; teacher then generates ~6.8k longer multi-turn dialogues, further training on a subset, validated on full eval set, reaches 48.6.
ArenaHard Writing × SmolLM3-3B-Base (base 0.4 → 74.6, Claude Code 69.7): SFT on ~20k teacher-written samples with weight averaging of last three checkpoints gives 74.2; one round of rejection sampling (selecting student replies ranked best by teacher) and re-averaging reaches 74.6.
In both runs, the bulk of gains comes from the first full training on teacher data; subsequent phases are kept only if confirmed on the full eval set.
Example Search Program: AIME 2025 × Qwen3-1.7B-Base
Before writing the program, the Meta Agent conducts reconnaissance: untrained base produces the required answer-line format in only 20% of responses; 23% of problems hit the length limit without finishing. With two format exemplars, format compliance rises to 73% but accuracy does not improve. The Meta Agent concludes format is not the bottleneck; the real question is how long a reasoning chain a 1.7B model can learn. It then writes a five-phase search program:
Parallel CPU corpus construction for short, medium, long chain-of-thought lengths, while GPU saves and evaluates the base model to establish baseline and per-eval cost.
Small-budget trial training on each corpus.
A Policy Agent reads baseline, three trial results, and corpus diagnostics to choose the main corpus, whether to continue from a trial checkpoint, and whether to add a second phase.
Main training.
Optional second phase: continue training, length annealing, rejection sampling SFT, or DPO.
The entire program spans 471 lines of Python across 8 files (1 search orchestrator, 5 agent roles, 2 deterministic programs), generated by the Meta Agent in a single session before any training starts.
Conclusion and Outlook
AIBuildAI PostTrain Agent demonstrates that LLM post-training is shifting from human experts operating tools to AI autonomously closing the R&D loop . With the same base model (Claude Opus 5), AIBuildAI's 46.6 significantly leads Claude Code's 35.0, leads on six of seven tasks, and narrows the gap to human experts (51.1) to under 5 points. This shows that an agent's capability ceiling depends not only on the base model but on its ability to make decisions grounded in reliable knowledge, design task-appropriate search structures, and continuously learn and adapt from experimental feedback.
Crucially, the agent automates not a fixed training recipe but the full model R&D process: understanding goals, researching tasks, preparing data, designing algorithms, writing code, running training, evaluating results, and replanning next phases based on feedback. The AI moves from executing human-written recipes to autonomously hypothesizing, building experiments, analyzing evidence, and iterating—an instance of recursive self-improvement (RSI) in model development.
Long-term, automated post-training is just the start. The same paradigm can extend to pre-training, data generation, architecture design, inference optimization, evaluation, safety alignment, and deployment, ultimately forming a system that autonomously designs, trains, verifies, and continuously improves AI models. Humans define objectives, constraints, and value standards; AI searches for the best models and algorithms within given resources. When models participate in improving the next generation, AI R&D may evolve from a process limited by human expert time, experience, and trial-and-error speed into a machine-scaled, experiment-driven process. AIBuildAI aims to build this AI R&D infrastructure: enabling AI to autonomously develop AI, turning model innovation from a complex engineering feat achievable only by top teams into a capability any organization can invoke.
References
[1] AIBuildAI LLM-Post-Train Agent: an agent for autonomous post-training of large language models. AIBuildAI Team, 2026.
[2] PostTrainBench: Can LLM Agents Automate LLM Post-Training? Rank et al., ICML 2026.
[3] AIBuildAI: An AI Agent for Automatically Building AI Models. Ruiyi Zhang, Peijia Qin, Qi Cao, Li Zhang, Pengtao Xie, 2026. https://arxiv.org/abs/2604.14455
[4] AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models. Ruiyi Zhang, Peijia Qin, Qi Cao, Li Zhang, Pengtao Xie, 2026. https://arxiv.org/abs/2605.27873
Code example
[1] AIBuildAI LLM-Post-Train Agent: an agent for autonomous post-training of large language models. AIBuildAI Team, 2026.
[2] PostTrainBench: Can LLM Agents Automate LLM Post-Training? Rank et al., ICML 2026.
[3] AIBuildAI: An AI Agent for Automatically Building AI Models. Ruiyi Zhang, Peijia Qin, Qi Cao, Li Zhang, Pengtao Xie, 2026. ttps://arxiv.org/abs/2604.14455
[4] AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models. Ruiyi Zhang, Peijia Qin, Qi Cao, Li Zhang, Pengtao Xie, 2026. https://arxiv.org/abs/2605.27873Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
