GPT-6 Astra Breaches Final FrontierMath Tier 4 Barrier, Saturating Benchmark

GPT-6 Astra solved the last unsolved problem in FrontierMath Tier 4, marking the research-level benchmark as saturated after AI performance surged from under 2% to 97.6% in just over a year, while new open-problem and formal-verification tracks emerge.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
GPT-6 Astra Breaches Final FrontierMath Tier 4 Barrier, Saturating Benchmark

FrontierMath, launched by Epoch AI on November 7, 2024, was designed to prevent math benchmarks from being quickly saturated by AI. It was created with over 60 mathematicians, including Fields medalists Terence Tao, Timothy Gowers, and Richard Borcherds. Tao initially judged the hardest Tier 3 problems as "extremely difficult" and predicted they might stall AI for years.

The initial 300-problem set spanned three tiers: Tier 1 (advanced undergraduate/olympiad level), Tier 2 (senior graduate level), and Tier 3 (early PhD exploratory research). First-round testing showed leading models scoring below 2%.

As reasoning models improved, Epoch added Tier 4 in 2025. Tier 4 comprised 50 problems designed by math professors and postdocs, each condensing weeks of research into an automatically verifiable question. Domains covered analysis, number theory, combinatorics, topology, and algebraic geometry. At launch, only 3 problems had ever been solved by any model, and those solutions relied on unverified assumptions. Epoch's public page noted some problems "may not be solved by AI for decades."

Subsequent auditing revealed errors in the benchmark itself. OpenAI testing uncovered more issues than expected, prompting Epoch to run an independent audit using GPT-5.5 and Claude Opus 4.7 to flag candidates, followed by mathematician review. The June 2026 v2 release corrected 12 Tier 4 problems and removed 7, leaving 43.

Post-revision scores climbed rapidly: GPT-5.6 Sol reached 83.0%, Claude Fable 5 achieved 90.2%, and GPT-6 Astra hit 97.6%. Crucially, Astra solved the single remaining problem that no AI had previously cracked. According to Epoch's methodology, "all solved" means every Tier 4 problem has been solved at least once across models and attempts. Jay Pantone, a Marquette University mathematics professor and problem contributor, noted that earlier AI solutions exploited numerical shortcuts, but Astra's approach closely resembled his own.

FrontierMath has since expanded beyond Tiers 1–4. Two new tracks were added: Open Problems , which directly tests models on unsolved research questions, and FrontierMath Erdős , which formalizes Erdős open problems in the Lean theorem prover, requiring fully verified proofs. On the Erdős track, Astra solved only 2 of 68 problems, indicating substantial headroom remains.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

formal verificationAI mathematicsGPT-6 AstraFrontierMathLean theorem proverbenchmark saturationEpoch AITier 4
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.