AI Peer Review Tested: GPT-5.5 Has Highest Recall but Generates Noise

ChenhaoTan · x · 2026-08-14

Discussing the limitations of evaluating AI peer-review tools based on a new benchmark where 100 known errors were planted into psychology papers. The author notes that while zero-shot GPT-5.5 achieved the highest recall, it generated an astonishing 512 comments—about 60% more than other systems.

The author argues that verification quality cannot be reduced to automated metrics like recall alone, emphasizing the need for real user feedback. They highlight that their SAI system achieved higher recall while maintaining a higher hit rate than competitors.

Related event: Multi-model AI Peer Review Catches 93% of Paper Errors(2 posts)→

Original post →

More from Research

Research channel →