Opus 5 looks perfect on benchmarks, but users say real-world quality is inconsistent
yunta_tsai · x · 2026-07-28
A user says Opus 5 feels bipolar: benchmark results look perfect, but real-world usage is inconsistent.
Their takeaway is that benchmarking may be hitting a limit in signal-to-noise ratio — the scores no longer track practical experience well.
More from Models
- Claude meme turns evals into poetry, joking about 0.03 deceptive alignment — maxsloef · 2026-07-28
- Gemini Flash 3.6 is claimed to match Sol 5.6 quality at 70% lower cost — bindureddy · 2026-07-28
- Anthropic’s Opus 4.8 beats version 5 for non-coding work, despite weaker benchmarks — burkov · 2026-07-28
- Thinking Machines releases Inkling, a 975B open-weights multimodal model with 1M context — paraschopra · 2026-07-28
- Users Say Opus 5 Looks Better After Repeated Bug-Fix and Feature Requests — Rasmic · 2026-07-28
- A user says GPT-5.4, Opus 4.6, and Kimi k3 already cover most needs — haider1 · 2026-07-28