LLM judges of AI-scientist novelty flip on tiny prompt tweaks, AI2 finds

TuhinChakr · x · 2026-10-09

A team led by Noy Sternlicht at Allen AI tested how much you can trust LLM judges when assessing the novelty of ideas produced by AI scientists.

Key findings:

Related event: Study Finds LLM Judges Unreliable at Rating Idea Novelty(2 posts)→

Original post →

More from Research

Research channel →