Study: LLM judges of AI-scientist idea novelty are unreliable

MarioKrenn6240 · x · 2026-10-09

A retweeted post notes that judging how novel an AI scientist's ideas are is now usually delegated to an LLM judge. The authors tested how much those judges can be trusted and found: not much — tiny prompt tweaks swing their verdicts wildly, sometimes below a coin flip.

Original post →

More from Research

Research channel →