Benchmark shows Opus 5 Max burns 3x tokens with minimal gain over Medium
PMinervini · x · 2026-08-27
PMinervini highlights findings from a benchmark pairing local LLMs with agent harnesses across 16 software engineering tasks. The results indicate that Opus 5 Max consumes roughly three times as many tokens as Opus 5 Medium without significant performance improvements, making it approximately 3x more expensive for comparable results. The benchmark spans Python, PyTorch, JAX, C, C++, Rust, and SQL in a sandboxed environment.
Related event: Opus 5 Max burns 3x tokens of Medium for marginal gains(2 posts)→
More from coding & agent
- Indicator Go v2: Go trading toolkit with AI integration via MCP — Shruti_0810 · 2026-08-27
- Shared Agent Skill Libraries Propagate Malware, 41.8% Self-Poisoning Rate Found — omarsar0 · 2026-08-27
- Tip: Use STRONG_GO keyword to stop agents from seeking constant confirmation — GabGarrett · 2026-08-27
- Integrating Warp Factory run costs into Slack for model routing optimization — vikvang1 · 2026-08-27
- Weaviate 1.39 ships Boost API and MMR to GA, adds 4-bit RQ quantization — CShorten30 · 2026-08-27
- Apodex 1.1 released: New model family and FrontierAgent framework — wuqiao · 2026-08-27