Study finds LLM novelty judges are unstable: scores swing wildly with evaluation design choices

_akhaliq · x · 2026-10-09

The paper "Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation" (arXiv:2610.02022) systematically tests whether LLMs can reliably judge research idea novelty.

A caution for AI-scientist style work: don't treat LLM novelty scores as reliable signals.

Related event: Study: LLM Judges Are Unreliable for Assessing Idea Novelty(3 posts)→

Original post →

More from Research

Research channel →