LLM graders can't match human taste on style eval — maybe why the gap persists
maxsloef · x · 2026-09-27
maxsloef built the Namedrop eval hoping labs would hillclimb on it, but couldn't create an automated grader matching his taste: across many models, prompts, and page presentations, LLM graders consistently diverged from his judgment. He speculates this may be why the style-elicitation issue persists — models' stylistic quality is hard to evaluate automatically, making targeted optimization difficult.
Related event: Blogger runs Namedrop eval in 1.5 hours of focused attention(4 posts)→
More from Models
- ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without — geoffwolfe · 2026-09-27
- GLM-5.3 Flash Matches Claude at 1/429th the Price in a YouTube Script Benchmark — OnlyProggingForFun · 2026-09-27
- Frontier AI is now so cheap and abundant that subscriptions go barely used — intellectronica · 2026-09-27
- Karpathy: Claude Opus 4.5 beats GPT-5 Pro for interactive history learning — doodlestein · 2026-09-27
- ChatGPT-6 Astra cracks 85-year-old 1941 Enigma message in two days — luisdans · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27