AI Agents Replicate 168 ICML Papers: Only 7 Completely Hold Up
profjamesevans · x · 2026-07-23
A team from the University of Chicago's SAI Labs used AI agents to review and replicate all 168 oral papers from top AI conference ICML, finding that only 7 papers completely held up to verification.
The team emphasized that this isn't just about peer review getting harder with AI; it reveals that peer review has never been a sufficiently strong verifier to facilitate cumulative scientific assembly. Top-to-bottom verification exposes the fundamental challenges of distributed science.
The study focused on three questions:
- Cost: The median estimated cost to reproduce a top ML paper.
- Blind spots: What code replication can reveal that narrative-only reviewing misses.
- Alignment: How well the AI review system aligns with existing human reviews.
The project has open-sourced all verification code and results.
More from Research
- Raji says FAccT papers are heavily cited in NIST, FTC and DOJ AI policy docs — rajiinio · 2026-07-23
- Nature piece says LLMs can forecast social-science experiment outcomes — RobbWiller · 2026-07-23
- Nature paper finds LLMs can predict social-science experiment results — RobbWiller · 2026-07-23
- Researchers open a demo for forecasting social-science treatment effects — RobbWiller · 2026-07-23
- LLM-only pilots cost under $1 and rival ~230-person human pilots — RobbWiller · 2026-07-23
- LLM forecasts were about 2x too large and weaker on field experiments — RobbWiller · 2026-07-23