AI reviews training AI reviewers: study finds 'scientific-judgment collapse' and an open-source fix
Sy-Tuyen Ho · hf · 2026-09-21
A controlled study examines the recursive risk of AI peer review: as model-generated reviews enter public data and future training corpora, later reviewers learn from earlier models' judgments.
Setup: Starting from Llama 3.1 8B, the authors fine-tune a reviewer on official ICLR 2018–2023 reviews, then train four successors on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews.
Finding: Synthetic reviews compress rating distributions and reduce both same-paper and corpus-level semantic diversity — dubbed "scientific-judgment collapse."
Mitigation: TrustReviewer, an open-source LLM reviewing system, intervenes at training time (curated single-stage corpus) and at test time (paired activation steering) without further training or expert annotation.
More from AGI Musings
- EA Called a Distinctive Experiment in High-Agency Applied Philosophy with Few Guardrails — jessi_cata · 2026-09-21
- Meta Muse pitch revives the question: will personal AI agents get real permission controls or just one big Allow button? — yi111 · 2026-09-21
- ctjlewis: nothing can be done now to stop AGI, like stopping the sunrise — ctjlewis · 2026-09-21
- Timothy B Lee: X-risk usage of RSI and superintelligence is borderline incoherent — binarybits · 2026-09-21
- Six principles for thinking about AI risk: the AI Snake Oil case against doom — binarybits · 2026-09-21
- AI Jam Founder Reflects: The Community Lost Sight of Its Builder Roots — Ghost_Pilot_MD · 2026-09-21