Planted 100 errors in papers to test AI peer review: best system caught 71, ensembles 93
alejandroll10 · x · 2026-09-03
Paul Litvak planted 100 known errors in 10 open-access psychology papers and ran them through frontier models and two commercial AI review tools, releasing all data and logs. Key findings:
- Best single system caught 71/100 errors; the worst caught 30.
- Pooling all systems caught 93/100 — model error detection is only partly correlated, making ensembling a big lever.
- Seven errors went undetected by every system; all were omissions (deleted information) rather than inserted mistakes.
- Refine.ink contributed the most unique catches but is expensive.
- False positives weren't measured; error distribution may not match real papers.
A replier notes their 'Forensic Metascience Agent' system, built in two weeks for stats auditing rather than peer review, scored 50 on the same benchmark.
More from Research
- SPAR to run RCTs testing whether secretly misaligned AI can sabotage human decisions — austinc3301 · 2026-09-03
- Materials scientist flags DiffCrysGen outputs violating charge neutrality, calls out garbage-rate reporting gap — CatAstro_Piyush · 2026-09-03
- Dynamical systems view of motor cortex: preparatory activity holds movement-specific code — burny_tech · 2026-09-03
- EleutherAI paper: persistent agent memory can enable 'authorization laundering' attacks — EleutherAI · 2026-09-03
- RealSWE benchmark: realistic user requests test coding agents, explicit intent boosts results — skku · 2026-09-03
- Six load forecasters benchmarked on GPU-hours: none beat the last-value baseline — Vegetable-Top-3670 · 2026-09-03