AI Agents Reproduce ICML Papers: Over 60% Fail, 99% Have Errors
量子位 · wechat · 2026-08-09
Recent studies reveal that AI agents are being used to systematically audit and reproduce top-tier conference papers, exposing a severe reproducibility crisis in academia.
- Reproduction Failure: An audit firm tested 168 ICML 2026 papers. Of the 92 with verifiable claims, only 34 could have over 40% of their conclusions successfully reproduced by AI, and just 8 reached over 80%. Failures stem from missing code, broken dependencies, or offline models.
- Rampant Errors: A GPT-5-based checker found that 99.2% of top AI papers contain at least one objective error (e.g., math mistakes), with the average number of errors per paper significantly increasing in recent years.
- AI Re-evaluating Science: Beyond catching modern errors, AI has identified flaws in century-old authoritative chemistry data. This marks the first time humanity has the technical capability to systematically review historical scientific literature at scale.
More from Research
- DeepMind's WeatherNext Model Published in Nature Gives 24h Extra Cyclone Lead Time — Scobleizer · 2026-08-09
- New Research: Optimizing Contact Timing for Fluid Humanoid Running and Gymnastics — Scobleizer · 2026-08-09
- AnyStyle Restyles 3D Scenes in 0.1s via Single-Pass, Accepted by ACM MM — Scobleizer · 2026-08-09
- China Builds Billion-Agent Society Simulator with Terrifying Predictive Power — pbaylies · 2026-08-09
- Leg-KILO: Open-Sourcing Robust SLAM for Dynamic Legged Robots — rsasaki0109 · 2026-08-09
- Physicist Argues AI Training Should Incentivize 'Hard-to-Vary' Explanations — shyamalanadkat · 2026-08-09