LLM graders can't match human taste on style eval — maybe why the gap persists

maxsloef · x · 2026-09-27

maxsloef built the Namedrop eval hoping labs would hillclimb on it, but couldn't create an automated grader matching his taste: across many models, prompts, and page presentations, LLM graders consistently diverged from his judgment. He speculates this may be why the style-elicitation issue persists — models' stylistic quality is hard to evaluate automatically, making targeted optimization difficult.

Related event: Blogger runs Namedrop eval in 1.5 hours of focused attention(4 posts)→

Original post →

More from Models

Models channel →