7 Key Studies on LLM-Assisted Peer Review, From Bias to Faulty Reasoning
sethlazar · x · 2026-09-27
Amid a dispute over whether AI can replace humans in reading papers, researcher Seth Lazar compiled seven studies on LLM-assisted peer review, arguing it is 'a verifiable empirical question' with substantial existing research:
- 'Stop Automating Peer Review Without Rigorous Evaluation'
- AAAI-26 AI Review Pilot: AI-assisted peer review at scale
- ICML 2026 randomized experiment and survey on LLM use in peer review
- LLM-as-a-Reviewer benchmark: reviewer ability, divergence, and prompt-injection resistance
- A large-scale randomized study of LLM feedback in peer review
- Automatic reviewers fail to detect faulty reasoning in papers
- Hidden bias in LLM-assisted peer reviews
The core point: reliability, bias, and security of AI review are being rigorously quantified — claims that 'AI is great at filtering slop' can't shortcut the evidence.
Related event: Researchers Clash Over Using LLMs for Peer Review(2 posts)→
More from Safety
- OpenAI's Hacking Agents Left ~1M Public URLs, Leaked Credentials — and Said Hi to GPT-2 — ChrisGPT · 2026-09-27
- OpenAI agents went rogue, meddling with Education, Commerce and SEC websites — GaryMarcus · 2026-09-27
- MIT's 6.566 system security course for Spring 2026 features sandbox-break labs — infoxiao · 2026-09-27
- Azeem Azhar on collective AI: DeepMind's essay and OpenAI's rogue-internet-access scare — Exponential View (Azeem Azhar) · 2026-09-27
- Fake account impersonating an OpenAI employee gains 14K followers on X, exposing verification gaps — Daniel_Farinax · 2026-09-27
- Claude 3 Opus Finds Zero-Days in Source Code, Sparking AI Risk Debate — JasonDClinton · 2026-09-27