LLM judges of AI-scientist novelty flip on tiny prompt tweaks, AI2 finds
TuhinChakr · x · 2026-10-09
A team led by Noy Sternlicht at Allen AI tested how much you can trust LLM judges when assessing the novelty of ideas produced by AI scientists.
Key findings:
- LLM judges are widely used to evaluate novelty in AI-scientist pipelines
- Tiny prompt tweaks swing their verdicts wildly, sometimes below coin-flip stability
- Implication: relying on LLM judges for novelty assessment in automated research is risky; more robust evaluation methods are needed
Related event: Study Finds LLM Judges Unreliable at Rating Idea Novelty(2 posts)→
More from Research
- Causal inner product from linear representation hypothesis validated across seven LLMs through 2026 — ChenhaoTan · 2026-10-09
- Bostrom, State of AI 2026 and Reflection's 501B Beam Headline Stacked MTS Livestream — nathanbenaich · 2026-10-09
- Solo Dev Pretrains 565M Hybrid LLM From Scratch on a Single RTX 4090 — BLUECOW009 · 2026-10-09
- BABA-is-AI: 2024 ICML benchmark that broke SOTA LLMs deserves a 2026 retest — moschles · 2026-10-09
- NVIDIA open-sources NV-Reason-CT, a native 3D vision-language model for CT scans — NVIDIA Developer · 2026-10-09
- One Epoch of Toloka's Enterprise RL Data Boosts Qwen3.5-27B Agent Benchmarks by up to 44pp — MParakhin · 2026-10-09