Planted 100 errors in papers to test AI peer review: best system caught 71, ensembles 93

alejandroll10 · x · 2026-09-03

Paul Litvak planted 100 known errors in 10 open-access psychology papers and ran them through frontier models and two commercial AI review tools, releasing all data and logs. Key findings:

A replier notes their 'Forensic Metascience Agent' system, built in two weeks for stats auditing rather than peer review, scored 50 on the same benchmark.

Original post →

More from Research

Research channel →