Claude Opus 5 tops every major benchmark but bottoms out on user experience

gerardsans · x · 2026-09-07

Responding to Nate Silver's critique of lab benchmarks, gerardsans argues AI may have hit a capability ceiling: labs keep leaning on flashy demos and evals for their press-release cycles, but these don't always translate into real value for consumers or businesses. He cites Claude Opus 5 — top ranking at every major eval yet absolute bottom in user feedback — as the poster child of the disconnect.

Original post →

More from AGI Musings

AGI Musings channel →