RECLAIM preprint: best AI agent reproduces only 41% of papers with code, 15% without
VraserX · x · 2026-09-29
Citing the new RECLAIM preprint, the author argues that an AI scientist worth paying for would spend most of its time checking other people's results.
- In the tier of papers with code, data, and weights available, the best agent reproduced only 41% of results
- Without code, the best reproduction rate dropped to 15%
- A common failure mode: agents implement the method as described without checking the numbers reported in the paper
The author concludes he would happily fund an AI lab that mostly replicates existing research — verification is exactly where current AI science capability falls short.
More from AGI Musings
- "If You Told 2020 That AI in 2026 Solved a Millennium Problem, You'd Call It the Singularity" — aran_nayebi · 2026-09-29
- Quintin Pope: cheap finetuning lets AIs defect, making durable AI coordination—and takeover—unlikely — QuintinPope5 · 2026-09-29
- When EA Ambition Means Buying Galaxies and Digital Descendants — abhiadesai · 2026-09-29
- Miles Brundage: not taking intelligence explosion seriously was my biggest recent mistake — Miles_Brundage · 2026-09-29
- Anthropic revenue chart lands in Museum of Chart Crimes for distorted time axis — infoxiao · 2026-09-29
- Dev claims 'AI brain rot' is real: the more he uses AI, the more he needs physical books — alexcovo_eth · 2026-09-29