NovGauge benchmark finds LLMs judge paper novelty poorly, hallucination rates up to 39%
sethlazar · x · 2026-09-14
Researchers introduce NovGauge, a fine-grained benchmark for diagnosing how well LLMs assess paper novelty — a key weak point as AI peer review gains traction.
- Built from 619 paper pairs and 50 multi-paper sets, sourced from ICLR reviewer novelty-overlap claims and survey co-citations, labeled along task, problem, and method dimensions.
- A cascading pipeline verifies per-dimension correctness, evidence grounding, and logical support.
- Across 18 LLMs, hallucination rates range from 0% to 39%; even among correct-positive judgments, over 70% cite evidence that fails to logically support the stated reason.
- Best model GPT-5.5 reaches only 43–72% Verified F1, and most models keep less than half their raw F1 after faithfulness checks.
Takeaway: even frontier models remain unreliable at judging novelty, arguing against delegating that part of peer review to AI for now.
More from Models
- Some $20 ChatGPT users report no 5-hour limit, only weekly caps — noletovictor · 2026-09-14
- Same 3D simulation prompt: Agnes 2.5 Pro Beta costs ~$0.20 vs ~$1.70 on GPT-5.6 Sol — iamaliveix · 2026-09-14
- Gary Marcus: GPT-6 Astra's costly gains don't fix OpenAI's broken economics — GaryMarcus · 2026-09-14
- Which 5 agent benchmarks actually matter? Ranking GPT-6 Astra vs Claude Fable 5.1 — IndyDevDan · 2026-09-14
- Mercury 2.5 reads a clinical chart, catches 3 planted errors in ~13 seconds — MaziyarPanahi · 2026-09-14
- OpenAI confirms Custom GPTs retirement, recommends Plugins as replacement — Longjumping_Log9999 · 2026-09-14