Dev complains Claude feels weird despite great evals, blames post-training quirks

MaziyarPanahi · x · 2026-09-14

A developer argues Claude Opus 5 and other Anthropic models score well on evals but feel strange in real use — over-checking files and work like they need babysitting. Citing a Dwarkesh podcast moment (23:25) where speakers claim Sonnet/Opus underperform GLM and Kimi, he suspects Anthropic is "training the weirdness in" during post-training, in ways benchmarks never catch.

Related event: Developers question why Claude benchmarks well but feels quirky in use(2 posts)→

Original post →

More from Models

Models channel →