Are Top Conference Papers Losing Credibility? AutoResearch Turns the Lens on Research Quality
An AI‑driven review of 168 ICML 2026 oral papers reveals that only 105 could be fully reproduced, with a median replication cost of $8,900, many hidden flaws, and 903 blind‑spot issues that human reviewers missed, questioning the trustworthiness of top‑conference publications.
Agent‑Driven Re‑evaluation of Top Conference Papers
Traditional peer review relies mainly on the paper text and does not require reviewers to set up environments or run code. The SAI Review system extracts key claims from each paper, obtains publicly released code, data, and models, executes the experiments, and compares the results claim‑by‑claim with the original manuscript.
Scope of the Study
The team examined all 168 oral papers accepted at ICML 2026. Of these, 105 papers were selected for full reproduction based on initial resource estimates, progressing from low‑resource to higher‑resource papers.
Reproduction Outcomes
Among the 105 fully reproduced papers:
67 papers could run the authors' released code directly.
38 papers required the system to re‑implement the experiments because no runnable code was provided.
Only 8 papers (≈8 %) achieved a reproduction score above 80 %.
For the 92 papers containing at least five verifiable claims, 34 papers scored above 40 % and 8 papers exceeded 80 %.
Cost of Full Replication
Using Google Cloud on‑demand pricing, the median cost to fully replicate all experiments of an ICML oral paper is about $8,900 . Seventeen papers cost over $100,000, with the most expensive approaching $2.2 million . These estimates include GPU time, storage, and compute for all experiments (including ablations and hyper‑parameter searches) but exclude personnel salaries and infrastructure overhead, which could multiply the total cost by 2–3×.
Common Reproducibility Obstacles
Among the 105 papers, 101 encountered at least one problem:
58 papers had code that could not run as released.
48 papers showed numerical discrepancies with the reported results.
42 papers lacked required data.
38 papers did not provide any runnable code.
4 papers depended on retired or unavailable models, making exact replication impossible.
Low reproduction scores do not necessarily imply intentional deception; factors such as code decay, proprietary datasets, and environment differences also contribute.
Blind‑Spot Analysis
SAI Review identified 903 issues that were not mentioned by human reviewers. Conversely, only 22 issues raised by humans were missed by the AI system. While the AI uncovered many code‑related problems, human reviewers still held an advantage in assessing novelty and research positioning.
Key Takeaways
Open‑source code does not guarantee that results are reproducible, and successful execution does not ensure that the original numbers can be regenerated. The study highlights a substantial gap between conference acceptance and verifiable scientific credibility, suggesting that the community should adopt more rigorous, execution‑based validation practices.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
