LLM-judged writing evals are blind to readability issues from reward model optimization
koltregaskes · x · 2026-09-07
sampaech cautions that any LLM-judged writing eval should be taken with a grain of salt: read the samples and judge for yourself. He argues LLM judges are largely blind to the readability problems that arise from optimizing on reward model preferences — a growing problem in the field.
More from Models
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07
- Astra looks great in demos but falls short on multi-step research-level reasoning, tester says — xiaosun86 · 2026-09-07
- GPT-6 Astra generates missile evasion simulation, showcasing stunning capability — algo_diver · 2026-09-07
- Hands-on: Astra demos are impressive but still fumbles multi-step research reasoning — xiaosun86 · 2026-09-07
- GPT 6's Blender skills reportedly trained by sub-Saharan artists paid $5/hour — Firm-Ad-2446 · 2026-09-07
- Model tracker Atlas ditches 'endangered' labels as newest Gemini may be riskier than Sonnet 3 — repligate · 2026-09-07