Calibrating Top-Tier Peer Review Scores with LLMs

新智元 · wechat · 2026-07-19

This article presents a study on peer reviews at top-tier conferences: as submissions to NeurIPS and ICLR continue to surge, reviewers' scoring criteria for identical comments have become increasingly inconsistent, leading to more random and unfair outcomes.

Instead of having LLMs directly replace reviewers, the research team proposes integrating a Calibration Layer in the background. This layer first structures comments into strengths and weaknesses, then uses an LLM to generate an AnchorScore based on unified standards as a reference. It then compares the deviation between the reviewer's original score and this reference. If the deviation is too large, it triggers a lightweight re-review, prompting the reviewer to provide additional justification or reconfirm their score.

Based on public data from ICLR 2023–2025, covering over 22,000 papers and 52,000 review records, they found that scoring scale drift is worsening year by year. Offline experiments also showed that within the hardest-to-judge boundary range of 5–6 points, the average citation count for papers filtered through this calibration was 94% higher than those filtered solely by manual scores.

Original post →

More from Research

Research channel →