Experiments show smarter models and higher effort write better LLM-judge evals

danshipper · x · 2026-10-08

An ongoing experiment on automated evals asks whether more compute means better checks from LLM judges — and the answer is 'yes, kind of.' Smarter models wrote better evals, and raising reasoning effort notably improved eval quality for small models like Luna.

Original post →

More from Research

Research channel →