Opus 5 is claimed to jump ahead on three spreadsheet-agent benchmarks
surmenok · x · 2026-07-25
Anthropic’s Opus 5 reportedly jumps ahead on spreadsheet-agent benchmarks
The post claims that Opus 5 is a major step up from Fable 5 and GPT-5.6 Sol on spreadsheet tasks.
What the chart shows
- Across three spreadsheet-agent benchmarks — Stage Gate, ShortcutBench V2, and ShortcutBench V1 — Opus 5 is plotted at the top of the accuracy frontier.
- The image claims it reaches 90.4% on Stage Gate, 78.1% on ShortcutBench V2, and 85.2% on ShortcutBench V1.
- It also argues Opus 5 matches or beats rivals at similar or lower wall-clock time and normalized cost.
- The caption says Fable 5 covers 55/57 ShortcutBench V2 tasks and 19/20 Stage Gate tasks at medium effort.
The overall message is that Opus 5 looks like another “step-function” leap for knowledge-work style spreadsheet agents.
More from Models
- Inclusion AI launches LLaDA2.2-flash, a diffusion model for agentic workloads — heyshrutimishra · 2026-07-25
- Opus 5 is said to discuss honesty 6× more than other agents in Village — bronzeagepapi · 2026-07-25
- Opus 5 reaches 30.2% on ARC-AGI 3 as critics question the benchmark — ChrSzegedy · 2026-07-25
- Opus 5 is catching bugs introduced by Opus 4.8 — damnGruz · 2026-07-25
- AMD open-sources Instella 16B MoE with checkpoints from pretraining to RL — bronzeagepapi · 2026-07-25
- Google posts a 1-hour agentic engineering course covering memory, MCP and multi-agent systems — ifioknkem · 2026-07-25