Is Mathematics Facing a Century‑Long Crisis? Tao Xue Zhan Warns of AI‑Driven Proof Overload
In his ICM 2026 talk, Fields Medalist Tao Xue Zhan argues that AI is turning mathematical proofs from scarce, high‑value achievements into abundant, hard‑to‑digest commodities, and he outlines a four‑stage roadmap and concrete safeguards for the community.
AI performance on research‑level mathematics
The independent First Proof benchmark released a second batch of ten open problems on 28 May 2024. Under controlled conditions four AI systems were evaluated by human experts. Seven of the ten problems were solved at publishable quality, with per‑problem compute costs ranging from $10 to $1,000.
Most evidence about AI’s mathematical ability is collected under biased, unscientific conditions; key cost data are often omitted.
Working assumption for conditional analysis
Temporarily assume that AI will soon be able to complete a substantial proportion of research‑level mathematical tasks at reasonable cost and quality. This assumption is not presented as a belief or endorsement, but as a basis for exploring consequences.
Traditional purposes of mathematical research
Solving open problems
Developing new theory
Understanding the world
Building a scholarly community
Training the next generation
Creating work of aesthetic value
Goodhart’s law in mathematics
When a metric becomes a target (e.g., “first to solve”), it ceases to be a good metric.
Evolving objectives as AI matures
Maximize the number of solved open problems.
Solve problems and verify correctness (e.g., with proof assistants such as Lean).
Solve, verify, and communicate results clearly.
In addition, achieve community digestion, acceptance, and canonicalization of results.
Risks identified
AI‑generated proofs may be correct but unreadable, hindering human understanding.
Many AI‑produced proofs submitted to the Erdős problems site lack human verification; some submitters explicitly state they lack the expertise to validate them.
Accumulation of proofs faster than the community can validate, write, peer‑review, and canonize – a phenomenon termed “proof digestion disorder.”
If we have a verified major theorem that no human can explain, the result is essentially useless.
Proposed mitigations
Mandatory disclosure of AI assistance in papers.
Reduce the prestige of “first to solve” and increase the value of “digestion” activities such as clear exposition, peer review, and canonicalization.
Hold authors fully responsible for correctness and citations when AI is used.
Restrict AI use in education and talent‑training contexts while allowing the mathematical community to define its own rules rather than ceding them to commercial AI incentives.
References
First Proof benchmark: https://1stproof.org/
Leiden AI & Mathematics Declaration: https://leidendeclaration.ai/
Erdős problems site: https://www.erdosproblems.com/
Code example
[1]https://1stproof.org/
[2]https://leidendeclaration.ai/
[3]https://www.erdosproblems.com/Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
