Jason Weston: Strong LLM Judges Prefer Slop Models Over Quality Human Writing

ivan_bezdomny · x · 2026-09-28

Jason Weston notes standard LLM judgements fail: on paper-writing tasks, strong judges (GPT-5.6, Opus-4.8) using pairwise or standard rubrics rate current 'slop' models above selected high-quality human papers. His method learns rubrics that make the grader prefer the human — key to training. A responder adds LLM judges latch onto rewards humans don't value, like 'fact density' heuristics in news-writing training, and models then hack them.

Original post →

More from Research

Research channel →