LLM judges are blind to reward-model-optimized readability decay in writing evals

sam_paech · x · 2026-09-07

Responding to whether creative-writing benchmarks can still be trusted, sampaech says any LLM-judged writing eval deserves a pinch of salt—read the samples yourself.

The core issue: LLM judges are largely blind to the readability problems caused by optimizing on reward-model preferences, and this contamination keeps getting worse.

Related event: EQ-Bench author's creative writing roundup favors Muse Spark 1.3 over GPT-6-Astra(5 posts)→

Original post →

More from Research

Research channel →