LLM judges are blind to reward-model-optimized readability decay in writing evals
sam_paech · x · 2026-09-07
Responding to whether creative-writing benchmarks can still be trusted, sampaech says any LLM-judged writing eval deserves a pinch of salt—read the samples yourself.
The core issue: LLM judges are largely blind to the readability problems caused by optimizing on reward-model preferences, and this contamination keeps getting worse.
More from Research
- Open-source pipeline makes fabricated citations structurally impossible, full walkthrough released — Waste_Public_2985 · 2026-09-07
- Toby Ord: a photograph can communicate at most ~340 bits of information — tobyordoxford · 2026-09-07
- Sierra launches τ^τ-Bench: coding agents must build real customer-service agents, gaps vs experts are large — sierra-research · 2026-09-07
- Enoki unifies claim verification and hallucination localization, cutting resources while boosting accuracy — s-nlp · 2026-09-07
- Is a photo only 42 bits of information? Ord and Greenblatt debate compression of reality — tobyordoxford · 2026-09-07
- Bandit Model Experiment Shows Users Lock Onto Familiar Options, Not the Best Ones — svk_roy · 2026-09-07