Study Finds LLM Judges Unreliable at Rating Idea Novelty

An Allen AI study finds that using LLM judges to assess the novelty of AI-generated research ideas is unreliable, as tiny prompt changes can drastically flip the verdicts.

2026-10-09 ~ 2026-10-09 · 2 related posts