Leaked Opus 5 Benchmarks: Major Jumps in Research Math and Long-Context

echen · x · 2026-07-30

Alleged benchmarks for the unreleased Opus 5 model have surfaced. Compared to Opus 4.8, the new version shows massive leaps in research math (Riemann-bench: 68.0% vs 47.2%), chart understanding (Chartography: 27.3% vs 15.9%), and long-context agents (HANDBOOK.md: 32.3% vs 21.9%). It shows modest improvements or ties in enterprise complex instructions, document reasoning, and creative writing.

Related event: Leaked Benchmarks Show Opus 5 Performance Leap(2 posts)→

Original post →

More from Models

Models channel →