Is Mathematics Facing a Century‑Long Crisis? Tao Xue Zhan Warns of AI‑Driven Proof Overload

In his ICM 2026 talk, Fields Medalist Tao Xue Zhan argues that AI is turning mathematical proofs from scarce, high‑value achievements into abundant, hard‑to‑digest commodities, and he outlines a four‑stage roadmap and concrete safeguards for the community.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Is Mathematics Facing a Century‑Long Crisis? Tao Xue Zhan Warns of AI‑Driven Proof Overload

AI performance on research‑level mathematics

The independent First Proof benchmark released a second batch of ten open problems on 28 May 2024. Under controlled conditions four AI systems were evaluated by human experts. Seven of the ten problems were solved at publishable quality, with per‑problem compute costs ranging from $10 to $1,000.

Most evidence about AI’s mathematical ability is collected under biased, unscientific conditions; key cost data are often omitted.

Working assumption for conditional analysis

Temporarily assume that AI will soon be able to complete a substantial proportion of research‑level mathematical tasks at reasonable cost and quality. This assumption is not presented as a belief or endorsement, but as a basis for exploring consequences.

Traditional purposes of mathematical research

Solving open problems

Developing new theory

Understanding the world

Building a scholarly community

Training the next generation

Creating work of aesthetic value

Goodhart’s law in mathematics

When a metric becomes a target (e.g., “first to solve”), it ceases to be a good metric.

Evolving objectives as AI matures

Maximize the number of solved open problems.

Solve problems and verify correctness (e.g., with proof assistants such as Lean).

Solve, verify, and communicate results clearly.

In addition, achieve community digestion, acceptance, and canonicalization of results.

Risks identified

AI‑generated proofs may be correct but unreadable, hindering human understanding.

Many AI‑produced proofs submitted to the Erdős problems site lack human verification; some submitters explicitly state they lack the expertise to validate them.

Accumulation of proofs faster than the community can validate, write, peer‑review, and canonize – a phenomenon termed “proof digestion disorder.”

If we have a verified major theorem that no human can explain, the result is essentially useless.

Proposed mitigations

Mandatory disclosure of AI assistance in papers.

Reduce the prestige of “first to solve” and increase the value of “digestion” activities such as clear exposition, peer review, and canonicalization.

Hold authors fully responsible for correctness and citations when AI is used.

Restrict AI use in education and talent‑training contexts while allowing the mathematical community to define its own rules rather than ceding them to commercial AI incentives.

References

First Proof benchmark: https://1stproof.org/

Leiden AI & Mathematics Declaration: https://leidendeclaration.ai/

Erdős problems site: https://www.erdosproblems.com/

Code example

[1]https://1stproof.org/
[2]https://leidendeclaration.ai/
[3]https://www.erdosproblems.com/
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AIMathematicsAI EthicsGoodhart's LawResearch WorkflowProof VerificationFirst ProofTao Xue Zhan
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.