Claude Opus 5 appears to beat Fable 5 on most benchmarks at half the price
Yuchenj_UW · x · 2026-07-25
Claude Opus 5 claims strong benchmark gains at half the price
The post says Opus 5 beats Fable 5 on nearly every benchmark and looks like a major jump in coding and agentic capability, with better token efficiency.
The attached benchmark chart shows Opus 5 leading or remaining highly competitive on:
- Agentic terminal coding: 43.3%
- Knowledge work: 1861
- ARC-AGI-3: 30.2%
- Agentic search: 90.8%
- Computer use: 70.6%
- Business workflows: 26.0%
- Biology: 49.4% hard / 90.1% human solved
The author’s key takeaway is that the model appears significantly better for coding and agentic work while costing half the price of Fable 5.
Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11