Claude Opus 5 tops every major benchmark but bottoms out on user experience
gerardsans · x · 2026-09-07
Responding to Nate Silver's critique of lab benchmarks, gerardsans argues AI may have hit a capability ceiling: labs keep leaning on flashy demos and evals for their press-release cycles, but these don't always translate into real value for consumers or businesses. He cites Claude Opus 5 — top ranking at every major eval yet absolute bottom in user feedback — as the poster child of the disconnect.
More from AGI Musings
- The three brainworm schools of AI discourse: denialist, x-risk, and toolism — mimi10v3 · 2026-09-07
- AI safety predictions keep turning from doomer nonsense to routine reality — DavidSKrueger · 2026-09-07
- Developer Yacine: shockingly little of my life progress was blocked by intelligence — yacineMTB · 2026-09-07
- Gary Marcus pushes back on Jensen Huang's AGI claim: no evidence, no definitions — Gary Marcus · 2026-09-07
- Grok tally: AGI has been declared "here" at least 4 times since Nov 2025 — suchenzang · 2026-09-07
- Emad Mostaque doubles down: AI will write nearly all code by 2027, developers gone by 2028 — jasonkneen · 2026-09-07