Are Top AI Labs Stopping Paper Reading? Inside a Large‑Scale Reproducibility Audit
A massive reproducibility audit of all 168 ICML 2026 oral papers reveals that only a handful meet basic verification standards, with median reproducibility below 30% and rerunning costs soaring into thousands of dollars, exposing a gap between paper prestige and scientific reliability.
In July 2023, the SAI project—founded by University of Chicago associate professor Tan Chenhao—conducted a large‑scale reproducibility review of every ICML 2026 oral paper (168 papers, representing the top 0.7% of 23,918 submissions). The audit duplicated the authors' workflow: reading the paper, inspecting experimental design, downloading code, models, and data, configuring the environment, and running the experiments to compare results with the reported claims.
SAI managed to fully reproduce 105 papers; only 104 of those had publicly released code. Of the reproduced papers, merely 34 (≈32%) reproduced more than 40% of the claimed results, and only 8 papers exceeded an 80% reproduction rate. The median reproducibility score hovered between 28% and 30% regardless of whether a paper contained one, three, or five verifiable claims.
Common failure modes included code that would not run, missing files, incomplete documentation, broken dependencies, and results that diverged from the paper. Four papers relied on models that had already been taken offline, making re‑creation impossible even with unlimited resources.
Two concrete examples illustrate the discrepancy between reported claims and actual artifacts. One paper advertised that only 0.77% of model parameters were trained, yet the released checkpoint showed 6.31% of parameters were updated—an eight‑fold increase over the claimed figure. Another paper presented a reliability table scored by a “judge model,” but the open‑source repository contained no such model nor any script capable of reproducing the table numbers.
Cost analysis, based on Google Cloud on‑demand pricing, showed a median expense of $8,900 to fully rerun an ICML oral paper. Seventeen papers cost over $100,000, and the most expensive case approached $2.2 million . Papers that depend heavily on large‑scale compute are therefore far more difficult for independent verification.
The audit highlights a systemic issue: papers without code are effectively “unassailable,” while even open‑source papers may be too costly to verify, allowing low‑quality work to pass peer review, accrue citations, and become a credential for hiring, admissions, or lab positions. Conferences rarely reopen review when reproducibility problems surface, and retractions are uncommon.
Despite the trend of some industry labs claiming they no longer read papers, the analysis shows that academic publications remain a crucial gate‑keeping mechanism for researchers seeking jobs, PhDs, or faculty positions. The paradox is that while large labs may de‑emphasize papers internally, they continue to use paper count, venue prestige, and citation metrics to filter external candidates.
Overall, the SAI findings suggest that the credibility of top‑tier AI research is under threat due to insufficient reproducibility, high verification costs, and a misalignment between paper prestige and actual scientific rigor.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
