AI Peer Review Tested: Ensembling Models Catches 93% of Planted Errors

ChenhaoTan · x · 2026-08-14

An in-depth evaluation of AI peer review tools reveals a massive performance gap among single systems, with the best catching 71 out of 100 planted errors and the worst only 30. The study highlights that because models are only partly correlated in the errors they find, ensembling multiple model outputs is a massive lever, catching 93% of the planted errors. However, all systems failed to detect omission-type errors where information was deleted. Furthermore, the commercial tool Refine.ink contributed the most unique catches, albeit at a higher cost.

Related event: Multi-model AI Peer Review Catches 93% of Paper Errors(2 posts)→

Original post →

More from Research

Research channel →