Developers question why Claude benchmarks well but feels quirky in use

A developer's viral essay argues that despite strong benchmark scores, Claude models behave oddly in practice, requiring constant babysitting, and questions whether Anthropic trained these quirks into the models during post-training.

2026-09-13 ~ 2026-09-14 · 2 related posts