LLM-judged writing evals are blind to readability issues from reward model optimization

koltregaskes · x · 2026-09-07

sampaech cautions that any LLM-judged writing eval should be taken with a grain of salt: read the samples and judge for yourself. He argues LLM judges are largely blind to the readability problems that arise from optimizing on reward model preferences — a growing problem in the field.

Related event: EQ-Bench author's creative writing roundup favors Muse Spark 1.3 over GPT-6-Astra(5 posts)→

Original post →

More from Models

Models channel →