Claude Opus 5 benchmark table shows strong early results across agentic tasks

FinanceYF5 · x · 2026-07-27

The post says Claude Opus 5 has just been released and people are already using it to build impressive things.

The attached benchmark-style image compares Opus 5 against Fable 5, Opus 4.8 and GPT-5.6 Sol across agentic coding, knowledge work, reasoning, computer use, workflows, legal and health tasks. Opus 5 leads or is competitive in several areas, including 43.3% on Frontier-Bench agentic terminal coding, 1861 on GPVal-AA knowledge work, 30.2% on ARC-AGI-3, 90.8% on BrowseComp, 70.6% on OSWorld 2.0, 26.0% on AutomationBench, and 49.4% on BioMysteryBench hard tasks.

The image also shows strong human-solved rates on biology and other task suites, making the post a compact snapshot of early Opus 5 benchmark claims.

Original post →

More from Models

Models channel →