Study: LLM judges of AI-scientist idea novelty are unreliable
MarioKrenn6240 · x · 2026-10-09
A retweeted post notes that judging how novel an AI scientist's ideas are is now usually delegated to an LLM judge. The authors tested how much those judges can be trusted and found: not much — tiny prompt tweaks swing their verdicts wildly, sometimes below a coin flip.
More from Research
- Exa launches ATLAS benchmark: even priciest search agents miss ~1/3 of results — yoimnotkesku · 2026-10-09
- PPTBench: a new benchmark testing if coding agents can rebuild visuals into editable slides — jiqizhixin · 2026-10-09
- TIDE attributes diffusion outputs to training images in milliseconds — serrjoa · 2026-10-09
- DeepScholar-Bench at COLM 2026: benchmarking AI-generated research synthesis — mrdrozdov · 2026-10-09
- Frontier AI models beat human experts at earnings predictions for the first time — maithra_raghu · 2026-10-09
- Cell paper reconstructs cell fate map of the mouse embryo — anshulkundaje · 2026-10-09