Users compile list of Claude Opus 5 complaints: best on benchmarks, worst to work with
gerardsans · x · 2026-08-20
A viral thread by @Ananth7e compiles consistent user complaints about Claude Opus 5: extremely verbose, overly apologetic and hedging, argues with instructions, invents weird jargon, does far more than asked, introduces subtle bugs that pass tests, confidently wrong then messes up fixes, and feels worse than Opus 4.8 for daily workflows. It remains strong at frontend, UI, and 3D work. The takeaway: best on benchmarks, worst to actually work with. Quoting the thread, @gerardsans argues this is the end of the "just scale more" game — benchmarks soar while daily usability collapses and users cancel.
Related event: Users Rebel Against Claude Opus 5: Benchmarks Up, Usability Down(4 posts)→
More from Models
- Grok 4.6 leads in legal/GDP benchmarks, lags in coding — ChrisGPT · 2026-08-20
- Leaked System Prompt: Domestic Giant's Client Uses 3-Layer Memory & MCP Routing — vista8 · 2026-08-20
- Ant Group opensources Ling-3.0 Base models with training checkpoints — 智东西 · 2026-08-20
- Test finds base model completions inherit AI/human writing traits from prompts — AaronBergman18 · 2026-08-20
- Polymarket: 10% chance a Chinese model tops the leaderboard by year-end — Polymarket · 2026-08-20
- User Reports Excessive GPT Refusals Blocking Tasks — cantrell · 2026-08-20