Benchmark shows Opus 5 Max burns 3x tokens with minimal gain over Medium

PMinervini · x · 2026-08-27

PMinervini highlights findings from a benchmark pairing local LLMs with agent harnesses across 16 software engineering tasks. The results indicate that Opus 5 Max consumes roughly three times as many tokens as Opus 5 Medium without significant performance improvements, making it approximately 3x more expensive for comparable results. The benchmark spans Python, PyTorch, JAX, C, C++, Rust, and SQL in a sandboxed environment.

Related event: Opus 5 Max burns 3x tokens of Medium for marginal gains(2 posts)→

Original post →

More from coding & agent

coding & agent channel →