Opus 5 is being compared to Opus 4.8 with a two-month gap and big benchmark jumps
SuhailKakar · x · 2026-07-25
The poster argues that Opus 5 is “twice as good” as Opus 4.8, noting only a two-month gap between the two releases and concluding that we may be closer to AGI than many people admit.
The attached chart compares benchmark performance across tasks such as agentic terminal coding, knowledge work, novel problem solving, and agentic search, with Opus 5 leading the older Opus 4.8 across the board shown in the image.
Related event: Anthropic Releases Claude Opus 5(57 posts)→
More from Models
- Qwen3.5-9B uncensored GGUF variant starts trending on Hugging Face — DavidAU · 2026-07-25
- Anthropic’s ARC-AGI-3 lead is being called meaningless as the benchmark saturates — morqon · 2026-07-25
- Opus 5 reportedly beats Fable 5 on long-horizon agent benchmarks, despite near-parity on single-shot tests — daniel_mac8 · 2026-07-25
- Claude Opus 5 is said to lead long-horizon coding and score 30.16% on ARC-AGI-3 — daniel_mac8 · 2026-07-25
- Opus 5 Adopts Frontier-Bench as Lead Benchmark One Day Post-Launch — ajratner · 2026-07-25
- Opus 5 debuts at No. 2 on Senior SWE-bench with 32% of Fable 5’s tokens — ajratner · 2026-07-25