Anthropic’s chart shows Claude Opus 5 leading several coding and knowledge benchmarks
Acceptable-Debt-294 · reddit · 2026-07-25
A Reddit post highlights Anthropic’s benchmark chart for Claude Opus 5, emphasizing its performance across coding, knowledge work, search, computer use, and biology.
Key numbers from the chart
- 43.3% on Frontier-Bench v1 (agentic terminal coding)
- 1861 on GDPval-AA v2 (knowledge work)
- 30.2% on ARC-AGI-3 (novel problem solving)
- 90.8% on BrowseComp (agentic search)
- 70.6% on OSWorld 2.0 (computer use)
- 68.8% on DeepScale v1.1 (agentic coding)
- 26.0% on AutomationBench (business workflows)
- 90.1% on BioMysteryBench (biology, human solved)
The chart also shows Opus 5’s cost-positioning versus Opus 4.8, Fable 5, and GPT-5.6 Sol, with Anthropic positioning it as a strong performance-per-dollar upgrade.
Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11