Can AI Do Independent Research? ASI‑Bench Measures Scientific Autonomy

ASI‑Bench, developed by Tsinghua and leading institutions, is a benchmark that evaluates AI’s scientific autonomy by progressively reducing method guidance across four levels, revealing that current models lose up to half their scientific score without detailed instructions, highlighting the gap to true independent research.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Can AI Do Independent Research? ASI‑Bench Measures Scientific Autonomy

Why ASI Matters

Recent AI progress relies on learning and recombining human‑known knowledge, achieving high scores on many exams and benchmarks. However, true scientific breakthroughs require exploring low‑probability, unknown spaces, formulating new hypotheses, and executing experiments without full human guidance. Measuring this capability is essential because without an evaluation, we cannot tell how far AI has progressed toward autonomous research.

What ASI‑Bench Is

ASI (Artificial Superintelligence) in this context is defined as an AI that moves from mastering existing knowledge to independently exploring unknown problems, creating new knowledge, and executing it. It requires three concurrent abilities:

General intelligence : applicability across diverse scientific domains.

Innovation : generating hypotheses and choosing methods when no complete solution is provided.

Autonomous execution : turning ideas into concrete steps, running experiments, debugging, and delivering verifiable results.

Existing benchmarks each test only one of these facets, leaving a gap for a comprehensive evaluation.

Benchmark Design

ASI‑Bench evaluates a single research project under four information‑gradient conditions (B1–B4):

B1 : Full method description and implementation steps – tests pure execution.

B2 : Only method name – tests ability to flesh out a method from its label.

B3 : Only research goal, raw data, and constraints – tests full scientific autonomy.

B4 : B3 plus irrelevant but correct distractors – tests robustness to misleading information.

The benchmark covers 60 project‑level tasks spanning 11 fields (mathematics, physics, chemistry, biology, astronomy, materials, earth science, medicine/biostatistics, computer science, robotics, electronic engineering). Each task undergoes a four‑eye, double‑review process, automated AI audit, and sandbox end‑to‑end execution to ensure reliable scoring.

Design core is information gradient
Design core is information gradient

Experimental Setup

Six agent harnesses (Codex, Apex Research, Claude Code, Kimi Code, MiMo Code, OpenHands) were paired with eight backbone models (GPT‑5.5, GPT‑5.6 Sol, Claude Opus, Kimi, DeepSeek, GLM, MiniMax, MiMo), forming 18 configurations. Each configuration was run on all 60 tasks.

Results

When full guidance (B1) was provided, the average Scientific Score across configurations was 50.92. Under minimal guidance (B3), the average dropped to 27.17, a loss of 23.75 points (≈46%). The largest drop occurred between B1 and B2 (average loss 21.28 points); the further drop from B2 to B3 was modest (≈2.5 points).

Top‑performing combos, such as Codex + GPT‑5.6 Sol, scored 71.78 at B1 and 51.60 at B3. Conversely, Claude Code + Claude Opus fell from 72.29 (B1) to 40.70 (B3). Failure analysis showed 62% of failures stemmed from scientific decision‑making (modeling, method selection, robustness judgment), while only ~26% were pure execution errors, indicating that the bottleneck is not arithmetic accuracy but the ability to decide “what to compute and why” after human guidance is removed.

Score distribution
Score distribution

Why the Numbers Matter

The benchmark’s low scores are intentional; they demonstrate a large discriminative space for future improvements. Each task’s construction involved expert curation, multi‑stage review, AI‑assisted auditing, and sandbox verification to ensure that high scores truly reflect scientific competence.

Who Should Use ASI‑Bench

Model teams : To prove that a new foundation or reasoning model is genuinely smarter, especially by improving B3 scores.

AI‑for‑Science developers : To diagnose whether their AI Scientist can continue research without full human instructions.

Agent framework developers : To compare different agent architectures on project‑level tasks that demand deeper reasoning.

Community Involvement

The first version’s 60 tasks were contributed by researchers from Tsinghua, MIT, CMU, USTC, Boston University, UIUC, Queensland, Flatiron Institute, Microsoft Research, and others. The benchmark invites further contributions of challenging, high‑value scientific problems, complete task specifications, scoring rubrics, and verification results.

MLNLP community logo
MLNLP community logo

Conclusion

ASI‑Bench provides the first clear measuring stick for AI’s progress toward autonomous scientific research. Current models show a substantial gap: they can execute tasks but struggle to independently design and steer research pathways. Closing this gap will be essential for moving beyond human‑level intelligence limits.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelsBenchmarkscientific researchagent evaluationAI autonomyASI-Bench
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.