Opus 5 benchmarks: #2 in research math and long-context agents, half the price
echen · x · 2026-07-30
Surge AI released Opus 5 benchmark results: Riemann-bench (research math) 68.0% (vs Opus 4.8 47.2%), Chartography (chart understanding) 27.3% (vs 15.9%), HANDBOOK.md (long-context agents) 32.3% (vs 21.9%). Anthropic positions it as near-fable intelligence at half the price. Overall not top 3, but strong in specific domains.
Related event: Leaked Benchmarks Show Opus 5 Performance Leap(2 posts)→
More from Models
- Dev tests Kimi K3: Full reasoning traces offer a transparent edge — doodlestein · 2026-07-30
- Dev builds parallel verification swarms leveraging cheap, fast Grok model — rudrank · 2026-07-30
- Glitch: Specific Prompts Cause Claude Opus to Leak Chain of Thought — matthen2 · 2026-07-30
- Researchers Find Anomalous Narrative Fulfillment Tendencies in Claude Opus 5 Base Mode — repligate · 2026-07-30
- Specific Prompt Triggers Anomalous User-Completion Behavior in Claude Opus 5 — matthen2 · 2026-07-30
- New Technologies Like MLA and GRPO are Decentralizing AI Open Source — zephyr_z9 · 2026-07-30