Jason Weston: Strong LLM Judges Prefer Slop Models Over Quality Human Writing
ivan_bezdomny · x · 2026-09-28
Jason Weston notes standard LLM judgements fail: on paper-writing tasks, strong judges (GPT-5.6, Opus-4.8) using pairwise or standard rubrics rate current 'slop' models above selected high-quality human papers. His method learns rubrics that make the grader prefer the human — key to training. A responder adds LLM judges latch onto rewards humans don't value, like 'fact density' heuristics in news-writing training, and models then hack them.
More from Research
- Stanford lands $25M NIH grant to build AI center for personalized dementia care — StanfordAILab · 2026-09-29
- RL Gains May Hit a Wall: Verifiable Data Limits and Stalling nanochart Benchmarks — QuintinPope5 · 2026-09-29
- TALES Benchmark on LLM Game-Playing Accepted to NeurIPS 2026, New Results Coming — tw_killian · 2026-09-29
- UBC's Schmidt wins ECML PKDD Test of Time Award for 2016 gradient convergence paper cited 2,000+ times — MarkSchmidtUBC · 2026-09-29
- Amazon AGI's AutoGym auto-generates tasks, environments, and verifiers for agent RL training — omarsar0 · 2026-09-29
- RECLAIM preprint: best AI agent reproduces only 41% of papers with code, 15% without — VraserX · 2026-09-29