AI reviews training AI reviewers: study finds 'scientific-judgment collapse' and an open-source fix

Sy-Tuyen Ho · hf · 2026-09-21

A controlled study examines the recursive risk of AI peer review: as model-generated reviews enter public data and future training corpora, later reviewers learn from earlier models' judgments.

Setup: Starting from Llama 3.1 8B, the authors fine-tune a reviewer on official ICLR 2018–2023 reviews, then train four successors on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews.

Finding: Synthetic reviews compress rating distributions and reduce both same-paper and corpus-level semantic diversity — dubbed "scientific-judgment collapse."

Mitigation: TrustReviewer, an open-source LLM reviewing system, intervenes at training time (curated single-stage corpus) and at test time (paired activation steering) without further training or expert annotation.

Original post →

More from AGI Musings

AGI Musings channel →