Developers question why Claude benchmarks well but feels quirky in use
A developer's viral essay argues that despite strong benchmark scores, Claude models behave oddly in practice, requiring constant babysitting, and questions whether Anthropic trained these quirks into the models during post-training.
2026-09-13 ~ 2026-09-14 · 2 related posts
- Long-read asks: is Anthropic training the weirdness into Claude? — MaziyarPanahi · 2026-09-13
- Dev complains Claude feels weird despite great evals, blames post-training quirks — MaziyarPanahi · 2026-09-14