Study Finds LLM Judges Unreliable at Rating Idea Novelty
An Allen AI study finds that using LLM judges to assess the novelty of AI-generated research ideas is unreliable, as tiny prompt changes can drastically flip the verdicts.
2026-10-09 ~ 2026-10-09 · 2 related posts
- Study: LLM judges of AI-scientist idea novelty are unreliable — MarioKrenn6240 · 2026-10-09
- LLM judges of AI-scientist novelty flip on tiny prompt tweaks, AI2 finds — TuhinChakr · 2026-10-09