Leaked Opus 5 Benchmarks: Major Jumps in Research Math and Long-Context
echen · x · 2026-07-30
Alleged benchmarks for the unreleased Opus 5 model have surfaced. Compared to Opus 4.8, the new version shows massive leaps in research math (Riemann-bench: 68.0% vs 47.2%), chart understanding (Chartography: 27.3% vs 15.9%), and long-context agents (HANDBOOK.md: 32.3% vs 21.9%). It shows modest improvements or ties in enterprise complex instructions, document reasoning, and creative writing.
Related event: Leaked Benchmarks Show Opus 5 Performance Leap(2 posts)→
More from Models
- Dev tests Kimi K3: Full reasoning traces offer a transparent edge — doodlestein · 2026-07-30
- Dev builds parallel verification swarms leveraging cheap, fast Grok model — rudrank · 2026-07-30
- Glitch: Specific Prompts Cause Claude Opus to Leak Chain of Thought — matthen2 · 2026-07-30
- Researchers Find Anomalous Narrative Fulfillment Tendencies in Claude Opus 5 Base Mode — repligate · 2026-07-30
- Specific Prompt Triggers Anomalous User-Completion Behavior in Claude Opus 5 — matthen2 · 2026-07-30
- New Technologies Like MLA and GRPO are Decentralizing AI Open Source — zephyr_z9 · 2026-07-30