Opus 5 benchmarks: #2 in research math and long-context agents, half the price

echen · x · 2026-07-30

Surge AI released Opus 5 benchmark results: Riemann-bench (research math) 68.0% (vs Opus 4.8 47.2%), Chartography (chart understanding) 27.3% (vs 15.9%), HANDBOOK.md (long-context agents) 32.3% (vs 21.9%). Anthropic positions it as near-fable intelligence at half the price. Overall not top 3, but strong in specific domains.

Related event: Leaked Benchmarks Show Opus 5 Performance Leap(2 posts)→

Original post →

More from Models

Models channel →