AI Peer Review Tested: Ensembling Models Catches 93% of Planted Errors
ChenhaoTan · x · 2026-08-14
An in-depth evaluation of AI peer review tools reveals a massive performance gap among single systems, with the best catching 71 out of 100 planted errors and the worst only 30. The study highlights that because models are only partly correlated in the errors they find, ensembling multiple model outputs is a massive lever, catching 93% of the planted errors. However, all systems failed to detect omission-type errors where information was deleted. Furthermore, the commercial tool Refine.ink contributed the most unique catches, albeit at a higher cost.
Related event: Multi-model AI Peer Review Catches 93% of Paper Errors(2 posts)→
More from Research
- WorldFM Open Source: Real-Time Multi-View Diffusion from Target Poses — tom_doerr · 2026-08-14
- Small Models Beat Large Ones in VLM Grounding with Tool Use — mervenoyann · 2026-08-14
- Google Introduces Promptable Gaze Target Estimation with Text and Visual Prompts — google · 2026-08-14
- Tensor-Level Quantization Boosts Gemma 4 12B Coding Performance by 8.5% — devildip · 2026-08-14
- DAB Benchmark: Simulating Messy Data Warehouses to Expose AI Agent Flaws — HamelHusain · 2026-08-14
- Parasma (YC): Human Brain Cells Achieve Next-Token Prediction — ycombinator · 2026-08-14