AI Agents Reproduce ICML Papers: Only 7 Out of 168 Fully Hold Up

A University of Chicago SAI Labs team utilized AI agents to comprehensively review and reproduce 168 ICML oral papers, finding that only 7 papers fully withstood verification. The experiment included 105 full reproductions. This stark contrast has sparked widespread discussion about the "reproducibility crisis" in the AI community and calls for academia to rethink its incentive mechanisms.

Confirmed

The research project (named SAI Review, based on OpenAIReview and Veritas) used AI agents to conduct reproduction tests on 168 ICML oral papers. The testing process included 105 full reproductions, and ultimately, only 7 papers largely "held up" or fully withstood scrutiny.

Why it matters

According to @profjamesevans, this initiative is not simply to prove that peer review has become harder in the AI era, but to reveal the current verification dilemma in academia. @ChenhaoTan points out that as AI agents make long-term research faster, academia should rethink its incentive mechanisms. Instead of merely pursuing a higher volume of papers, more emphasis should be placed on the verifiability and quality of the research. This event quantitatively and intuitively exposes the pre-existing "reproducibility crisis" in academia.

2026-07-23 ~ 2026-07-24 · 5 related posts

Primary sources