Anthropic’s chart shows Claude Opus 5 leading several coding and knowledge benchmarks
Acceptable-Debt-294 · reddit · 2026-07-25
A Reddit post highlights Anthropic’s benchmark chart for Claude Opus 5, emphasizing its performance across coding, knowledge work, search, computer use, and biology.
Key numbers from the chart
- 43.3% on Frontier-Bench v1 (agentic terminal coding)
- 1861 on GDPval-AA v2 (knowledge work)
- 30.2% on ARC-AGI-3 (novel problem solving)
- 90.8% on BrowseComp (agentic search)
- 70.6% on OSWorld 2.0 (computer use)
- 68.8% on DeepScale v1.1 (agentic coding)
- 26.0% on AutomationBench (business workflows)
- 90.1% on BioMysteryBench (biology, human solved)
The chart also shows Opus 5’s cost-positioning versus Opus 4.8, Fable 5, and GPT-5.6 Sol, with Anthropic positioning it as a strong performance-per-dollar upgrade.
Related event: Anthropic Releases Claude Opus 5(40 posts)→
More from Models
- Opus 5 appears to improve on ARC-AGI 1 and 2, and may rely on algebraic puzzle solving — herbiebradley · 2026-07-25
- Anthropic says Opus 5 is its hardest model to trick with prompt injection — cedric_chee · 2026-07-25
- Kimi K3 trails Mythos on cyber-range tasks in a benchmark chart shared online — petrusenko_max · 2026-07-25
- Polymarket puts U.S. AI safety bill odds at 34% as Claude Opus 5 surfaces — Polymarket · 2026-07-25
- Moonshot AI launches Kimi K3 with 2.8T parameters and a 1M-token context — dl_weekly · 2026-07-25
- Anthropic’s Claude releases appear to have sped up from every four months to monthly in 2026 — dustinvtran · 2026-07-25