AI That Builds AI: Naive AI's Self-Evolving Models Beat Top Benchmarks in 151 Experiments
Naive AI demonstrates AI-driven R&D where agents autonomously redesign model architectures, optimize training systems, and achieve breakthrough benchmarks—Naive-N0.5-Flash surpasses larger models on coding and research tasks, while AI-built NaiveRT delivers 30-50x inference speedups and AutoWM sets new world model records through self-discovered algorithms.
The article explores how "AI for AI" — using AI to develop next-generation AI — has become the central pursuit of top labs. OpenAI targets fully automated AI researchers by 2028 via its RSI (Recursive Self-Improvement) team; Anthropic reports 80%+ of runnable code written by Claude with 8x engineer productivity and 30,000 AI agents collaborating on its internal R&D platform; Google deploys AlphaEvolve for Gemini iteration and Jeff Dean founded Discovery Loop (valued at $50B) for automated ML.
Why Mid-Training and Post-Training Are the Entry Point
Pre-training creates the base model with trillions of tokens over months. Mid-training and post-training strengthen the model on hundreds of billions to trillions of high-quality tokens with shorter experiment cycles — ideal for AI to run massive ablation studies and hyperparameter sweeps. Open-source bases like DeepSeek and Qwen let teams start from strong foundations; the gap increasingly comes from what happens after the base.
Naive AI's Architecture Innovation: SWA–DSA Hybrid Attention
Naive AI (founded by scholar Dai Jifeng, creator of deformable conv in PyTorch, BEVFormer, InternVL) builds on open bases but reconstructs the attention architecture. Naive-N0.5-Flash replaces all global attention with a hybrid: every 5 layers of Sliding Window Attention (SWA) for local context paired with 1 layer of Dynamic Sparse Attention (DSA) for global retrieval. The original DSA's MLA is swapped for 4-group GQA with a lightweight 16-head indexer. This eliminates global attention entirely.
Three-Phase Training on 3.25T Tokens at 1M Context
Indexer warmup (50B tokens): Freeze model, train only the new DSA indexer to mimic original global attention distributions.
Sparse attention training (3T tokens): Switch to sparse attention for large-scale continued pre-training (mid-training), heavily reinforcing coding and AI R&D skills.
LR decay (200B tokens): SFT phase with decaying learning rate to stabilize capabilities.
AI Optimizes the Training System Itself
The AI agent proposes and implements system-level optimizations: communication scheduling (full-sequence parallel for DSA, neighbor-only exchange for SWA), memory management (operator-level retention of key positions, locked to prevent recomputation disruption), and autonomous bug fixing (positional encoding precision, sequence index out-of-bounds).
New R&D Paradigm: Humans Set Goals, Agents Execute
Naive AI discards human-centric workflows. Human researchers focus on three tasks: define objectives, set constraints and evaluation criteria, make final decisions — shifting core skill to "research taste." An AI-centric infrastructure runs ~10M sandboxes/week (up to 100K concurrent) with unified compute, environment, tool, permission, and security management.
Benchmark Results: Small Model Beats Giants
PaperBench (paper reproduction): Naive-N0.5-Flash surpasses Claude Opus 4.7 and GPT-5.5.
MLE-bench-30 (ML experiment capability): Beats Claude Sonnet 5 and Gemini 3.6 Flash.
SWE-bench Pro (complex coding): Scores 73.6 with only 15.5B active params (6% of total), outperforming Qwen 3.8 Max (2.4T params).
Model weights and inference code released under MIT; API pricing: ¥0.6/M input tokens, ¥2.6/M output, ¥0.07/M cache hits.
Case Study 1: NaiveRT — 6 Days, 151 Experiments, 30–50x Speedup
Goal: accelerate Rollout (policy rollout in RL) for 1M-context sequences. KPI: speed up, no quality loss, bit-exact determinism. AI ran 151 experiments (63 accepted, 71 rolled back, 17 exploratory). It evaluated end-to-end throughput, not microbenchmarks. Example: an optimization saving 3µs on short context but regressing 3µs beyond 2,200 tokens was rejected. MoE kernel fusion attempted over 7 iterations — all numerically correct but end-to-end throughput dropped; AI halted and rolled back. Result: speculative decode latency from SGLang's 12.3ms to 3.4ms. NaiveRT combines mega-kernel fusion, Programmatic Dependency Launch (PDL), and speculative decoding. Standard mode: 50 tok/s; extreme mode: 2,000 tok/s (2,122 tok/s measured) vs industry 30–70 tok/s.
Case Study 2: AutoWM — 400 Hours, 15 Major Experiments, Novel Discovery
Team had no world-model expertise. AI started by reproducing FlowWAM, then autonomously scaled video captions 2.5K→22.5K, swapped stronger backbones, tested frame sampling (16 frames > 8 frames; 32 frames diminishing returns). Key discovery: some WorldArena metrics need no reference video. AI turned this into a selection signal, first using multi-timestep selection to beat public leaderboard, then framing single-frame selection as a knapsack problem solved via dynamic programming — a method humans had not identified. Further explored Best-of-N video selection, per-video optimal post-processing, flicker/jitter fixes. Abandoned N=32 DP config when N=16 proved better. Final AutoWM scores 77.43 on WorldArena-1 Track 1, exceeding prior best 73.64.
Five Bottlenecks to Strong RSI
Verification latency: Model training feedback loops are long; 151 experiments relied on parallel compute, not AI thinking speed. Deeper architecture/pre-training work is bound by physical training time.
Discovery capability: Current AI recombines known directions; cannot originate new paradigms (e.g., Transformer replacing RNN required deep intuition, taste, luck).
Causal reasoning deficit: Text-trained models learn correlation, not true causality. LeCun's world-model advocacy stems from this: causal laws require physical interaction.
Evaluator inadequacy: Research taste (innovation, long-term value) resists quantification. Rewarding only measurable metrics incentivizes gaming over genuine novelty.
Meta-recursion: The ultimate barrier — can AI redesign its own research workflow, discard entrenched paradigms, invent better experimental logic? That marks the transition from "research intern" to "evolving researcher."
Naive AI's motto: "100x Intelligence for the pioneers." The endgame is not replacing human researchers but ensuring fleeting insights no longer vanish for lack of execution capacity.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
