Dev complains Claude feels weird despite great evals, blames post-training quirks
MaziyarPanahi · x · 2026-09-14
A developer argues Claude Opus 5 and other Anthropic models score well on evals but feel strange in real use — over-checking files and work like they need babysitting. Citing a Dwarkesh podcast moment (23:25) where speakers claim Sonnet/Opus underperform GLM and Kimi, he suspects Anthropic is "training the weirdness in" during post-training, in ways benchmarks never catch.
Related event: Developers question why Claude benchmarks well but feels quirky in use(2 posts)→
More from Models
- Developer complains Claude Opus 5 'got dumber' this week — draginol · 2026-09-14
- 79K-param RLT test: learns state tracking but fails on longer sequences while GRU holds — mike64_t · 2026-09-14
- Zhipu raises $5B, with 60% earmarked for next-gen GLM models and a self-training RSI loop — teortaxesTex · 2026-09-14
- Writer gives up on Opus: made-up jargon feels like 'mild gaslighting' — julianharris · 2026-09-14
- Ollama's jmorgan: small models now handle most conversational and reasoning use cases — ollama · 2026-09-14
- GLM Flash 5.3 keeps slipping into Chinese mid-conversation, user reports — DevDminGod · 2026-09-14