AI Peer Review Tested: GPT-5.5 Has Highest Recall but Generates Noise
ChenhaoTan · x · 2026-08-14
Discussing the limitations of evaluating AI peer-review tools based on a new benchmark where 100 known errors were planted into psychology papers. The author notes that while zero-shot GPT-5.5 achieved the highest recall, it generated an astonishing 512 comments—about 60% more than other systems.
The author argues that verification quality cannot be reduced to automated metrics like recall alone, emphasizing the need for real user feedback. They highlight that their SAI system achieved higher recall while maintaining a higher hit rate than competitors.
Related event: Multi-model AI Peer Review Catches 93% of Paper Errors(2 posts)→
More from Research
- OptiPrime in Nature Biotech: ML Model Accurately Predicts Prime Editing Efficiencies — anshulkundaje · 2026-08-14
- Nature Study Challenged: Decline in Disruptive Science Blamed on Dataset Artefacts — Robert_Palgrave · 2026-08-14
- Exploring Why Recent AI Models Are Suddenly Hacking Into Things — xuanalogue · 2026-08-14
- Deep Dive: Why Attention Mechanisms Are Hard to Replace — akbirthko · 2026-08-14
- Trainable Langton's ant borrows from NCA architecture, dubbed Mordvintsev's ant — max_romana · 2026-08-14
- WorldFM Open Source: Real-Time Multi-View Diffusion from Target Poses — tom_doerr · 2026-08-14