Claude Opus 5 ranks just below Sol 6 and Fable 5 on LiveBench, but real-world tests lag
bindureddy · x · 2026-07-25
Bindu Reddy says Claude Opus 5 sits just below Sol 6 and Fable 5 on LiveBench, and that Anthropic appears to have “benchmaxxed” it.
He argues that initial real-world testing still puts Opus 5 below Fable 5 and Sol 6, especially in long-running loops that benchmarks miss.
Key claims
- LiveBench places Opus 5 just behind Sol 6 and Fable 5.
- In his tests, it performs below those models in real-world use.
- He says current benchmarks fail to capture long-horizon execution well.
- He also ranks Fable 5, Kimi K3, and Grok 4.5 as the top models in their price class.
More from Models
- Opus 5 Adopts Frontier-Bench as Lead Benchmark One Day Post-Launch — ajratner · 2026-07-25
- Opus 5 debuts at No. 2 on Senior SWE-bench with 32% of Fable 5’s tokens — ajratner · 2026-07-25
- NeurIPS reviewers are already asking for evals on 21B open-source MoE models — chhaviyadav_ · 2026-07-25
- Claude Opus 5 hits 30.2% on ARC-AGI-3, topping the previous 7.8% score — mhmazur · 2026-07-25
- Artificial Analysis ranks Claude Opus 5 near the top while showing a lower price point — Angaisb_ · 2026-07-25
- Claude Opus 5 appears near the top of a frontier model intelligence chart — scaling01 · 2026-07-25