NovGauge benchmark finds LLMs judge paper novelty poorly, hallucination rates up to 39%

sethlazar · x · 2026-09-14

Researchers introduce NovGauge, a fine-grained benchmark for diagnosing how well LLMs assess paper novelty — a key weak point as AI peer review gains traction.

Takeaway: even frontier models remain unreliable at judging novelty, arguing against delegating that part of peer review to AI for now.

Original post →

More from Models

Models channel →