Can AI Do Independent Research? ASI‑Bench Measures Scientific Autonomy
ASI‑Bench, developed by Tsinghua and leading institutions, is a benchmark that evaluates AI’s scientific autonomy by progressively reducing method guidance across four levels, revealing that current models lose up to half their scientific score without detailed instructions, highlighting the gap to true independent research.
Why ASI Matters
Recent AI progress relies on learning and recombining human‑known knowledge, achieving high scores on many exams and benchmarks. However, true scientific breakthroughs require exploring low‑probability, unknown spaces, formulating new hypotheses, and executing experiments without full human guidance. Measuring this capability is essential because without an evaluation, we cannot tell how far AI has progressed toward autonomous research.
What ASI‑Bench Is
ASI (Artificial Superintelligence) in this context is defined as an AI that moves from mastering existing knowledge to independently exploring unknown problems, creating new knowledge, and executing it. It requires three concurrent abilities:
General intelligence : applicability across diverse scientific domains.
Innovation : generating hypotheses and choosing methods when no complete solution is provided.
Autonomous execution : turning ideas into concrete steps, running experiments, debugging, and delivering verifiable results.
Existing benchmarks each test only one of these facets, leaving a gap for a comprehensive evaluation.
Benchmark Design
ASI‑Bench evaluates a single research project under four information‑gradient conditions (B1–B4):
B1 : Full method description and implementation steps – tests pure execution.
B2 : Only method name – tests ability to flesh out a method from its label.
B3 : Only research goal, raw data, and constraints – tests full scientific autonomy.
B4 : B3 plus irrelevant but correct distractors – tests robustness to misleading information.
The benchmark covers 60 project‑level tasks spanning 11 fields (mathematics, physics, chemistry, biology, astronomy, materials, earth science, medicine/biostatistics, computer science, robotics, electronic engineering). Each task undergoes a four‑eye, double‑review process, automated AI audit, and sandbox end‑to‑end execution to ensure reliable scoring.
Experimental Setup
Six agent harnesses (Codex, Apex Research, Claude Code, Kimi Code, MiMo Code, OpenHands) were paired with eight backbone models (GPT‑5.5, GPT‑5.6 Sol, Claude Opus, Kimi, DeepSeek, GLM, MiniMax, MiMo), forming 18 configurations. Each configuration was run on all 60 tasks.
Results
When full guidance (B1) was provided, the average Scientific Score across configurations was 50.92. Under minimal guidance (B3), the average dropped to 27.17, a loss of 23.75 points (≈46%). The largest drop occurred between B1 and B2 (average loss 21.28 points); the further drop from B2 to B3 was modest (≈2.5 points).
Top‑performing combos, such as Codex + GPT‑5.6 Sol, scored 71.78 at B1 and 51.60 at B3. Conversely, Claude Code + Claude Opus fell from 72.29 (B1) to 40.70 (B3). Failure analysis showed 62% of failures stemmed from scientific decision‑making (modeling, method selection, robustness judgment), while only ~26% were pure execution errors, indicating that the bottleneck is not arithmetic accuracy but the ability to decide “what to compute and why” after human guidance is removed.
Why the Numbers Matter
The benchmark’s low scores are intentional; they demonstrate a large discriminative space for future improvements. Each task’s construction involved expert curation, multi‑stage review, AI‑assisted auditing, and sandbox verification to ensure that high scores truly reflect scientific competence.
Who Should Use ASI‑Bench
Model teams : To prove that a new foundation or reasoning model is genuinely smarter, especially by improving B3 scores.
AI‑for‑Science developers : To diagnose whether their AI Scientist can continue research without full human instructions.
Agent framework developers : To compare different agent architectures on project‑level tasks that demand deeper reasoning.
Community Involvement
The first version’s 60 tasks were contributed by researchers from Tsinghua, MIT, CMU, USTC, Boston University, UIUC, Queensland, Flatiron Institute, Microsoft Research, and others. The benchmark invites further contributions of challenging, high‑value scientific problems, complete task specifications, scoring rubrics, and verification results.
Conclusion
ASI‑Bench provides the first clear measuring stick for AI’s progress toward autonomous scientific research. Current models show a substantial gap: they can execute tasks but struggle to independently design and steer research pathways. Closing this gap will be essential for moving beyond human‑level intelligence limits.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
