Anthropic’s Opus 4.8 beats version 5 for non-coding work, despite weaker benchmarks
burkov · x · 2026-07-28
Anthropic’s Opus 4.8 is described as better than version 5 for non-coding work, even though Opus 5 allegedly beats Fable 5 on benchmarks.
The post argues that this looks like a “benchmaxed” model: stronger benchmark scores, but worse real-world usefulness than 4.8 on general tasks. It also says the author still prefers Codex for coding, and is hesitant about Kimi K3 because it appears slower and more expensive, which the post attributes to China lacking enough high-end GPUs.
Related event: Anthropic Opus 5 Leads Benchmarks but Splits Real-World Reviews(6 posts)→
More from Models
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11