Opus 5 ranks high on benchmarks but still feels slippery in practice
brandon_galang · x · 2026-07-29
The author says Opus 5 does not feel as clearly better in practice as its benchmark position suggests.
They stress two points:
- They do not think Opus 5 is a bad model.
- But for workflows that are not one-shot, the model’s behavior ergonomics matter more, and Opus 5 feels more elusive and slippery than Fable 5, which felt obviously better immediately.
The takeaway is that raw benchmark strength does not automatically translate into a better day-to-day model experience.
Related event: Opus 5 Scores High on Benchmarks but Underperforms in Practice(4 posts)→
More from Models
- Together AI and Moonshot AI set July 30 webinar on Kimi K3 architecture — togethercompute · 2026-07-29
- Three frontier models missed a simple link-extraction task and rewrote it as a web app — NickPassig · 2026-07-29
- Fable is being praised for trying parallel and multi-stream solutions first — dejavucoder · 2026-07-29
- Kimi K3’s intro feels unusually information-dense in very few tokens — nathanbenaich · 2026-07-29
- MoonshotAI’s FlashKDA commit hints at a hybrid attention mainline model — peterjliu · 2026-07-29
- User says Opus 5 ignores tools, breaks workflows, and feels overly ignorant — omarsar0 · 2026-07-29