AI agents replicated ICML 2026 oral papers, and only 7 mostly held up
ChenhaoTan · x · 2026-07-23
A repost of a research thread about whether science is actually verifiable. The authors replicated ICML 2026 oral papers with AI agents and report that only 27 of 105 papers reproduced more than 40% of the targeted conclusions, while only 7 reproduced more than 80%.
Key findings:
- AI review covered 78% of the comments on which two or more human reviewers agreed.
- For code and reproducibility, the AI reviewers found 903 issues that human reviewers missed, while humans found only 22 issues missed by the AI system.
- Humans still did better on novelty and positioning, suggesting that taste and research judgment remain hard to automate.
The broader argument is that agentic verification may become a useful complement to narrative-only peer review.
Related event: AI Agents Reproduce ICML Papers: Only 7 Out of 168 Fully Hold Up(5 posts)→
More from AGI Musings
- mark_k: "Eject all doomers from the AI companies — they're destroying you from the inside" — mark_k · 2026-09-11
- Adam Marblestone's Podcast Reading List: Evolution of Intelligence to Digital Minds — KordingLab · 2026-09-11
- Superintelligence will be maximum good, not stupid or evil, argues Patterson — davidpattersonx · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Should AI models be taught morality? Breakout incidents expose missing ethical training — Pfungus_ · 2026-09-11
- SoftBank's Masayoshi Son predicts 100 trillion self-replicating AIs: "humans' era as top life form is ending" — Puzzleheaded-King584 · 2026-09-11