Study finds LLM novelty judges are unstable: scores swing wildly with evaluation design choices
_akhaliq · x · 2026-10-09
The paper "Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation" (arXiv:2610.02022) systematically tests whether LLMs can reliably judge research idea novelty.
- Automated idea generation is increasingly judged by LLMs, but such judges are rarely validated — and when validated, it's on human-authored papers, not generated ideas
- The authors auto-built an evaluation set from OpenReview, mining passages where reviewers explicitly affirmed or disputed originality, keeping only unanimous submissions
- Result: the judges perform poorly, and novelty scores are highly sensitive to design choices
A caution for AI-scientist style work: don't treat LLM novelty scores as reliable signals.
Related event: Study: LLM Judges Are Unreliable for Assessing Idea Novelty(3 posts)→
More from Research
- China Telecom and MemTensor unveil HaluMem, first operation-level benchmark for agent memory hallucinations — jiqizhixin · 2026-10-09
- BAAI's AREX research agent checks answers requirement-by-requirement, hits 82.5% BrowseComp — DeepLearningAI · 2026-10-09
- New piece lays out how to build RL environments aimed at superintelligence — JenniferHli · 2026-10-09
- Quantum Counterfactuals: Quantum RNGs as an Exploration Source for RL — jessi_cata · 2026-10-09
- OpenAI theorem drop collides with researchers' work: stronger bounds but 'unreadable' proof — guyvdb · 2026-10-09
- AI-written science floods preprint servers; researchers propose decision language models as filter — lpachter · 2026-10-09