Claude Opus 5 benchmark table shows strong early results across agentic tasks
FinanceYF5 · x · 2026-07-27
The post says Claude Opus 5 has just been released and people are already using it to build impressive things.
The attached benchmark-style image compares Opus 5 against Fable 5, Opus 4.8 and GPT-5.6 Sol across agentic coding, knowledge work, reasoning, computer use, workflows, legal and health tasks. Opus 5 leads or is competitive in several areas, including 43.3% on Frontier-Bench agentic terminal coding, 1861 on GPVal-AA knowledge work, 30.2% on ARC-AGI-3, 90.8% on BrowseComp, 70.6% on OSWorld 2.0, 26.0% on AutomationBench, and 49.4% on BioMysteryBench hard tasks.
The image also shows strong human-solved rates on biology and other task suites, making the post a compact snapshot of early Opus 5 benchmark claims.
More from Models
- TamilLM finishes pre-training on 30B tokens, cutting held-out loss to 2.84 — sachinmaya1980 · 2026-07-27
- Several frontier models solve a stubborn distributed-systems problem with careful prompting — _xjdr · 2026-07-27
- Claude Opus 5 demo builds a full brand from one prompt — FinanceYF5 · 2026-07-27
- Reddit users joke that Gemini “died again” — hebittoken · 2026-07-27
- A repost says treating Opus 3 seriously is a kind of superpower — repligate · 2026-07-27
- User says Gemini Pro keeps erroring and failing Gmail Workspace tasks — CleanDifference6455 · 2026-07-27