Grok 4.5 tops several coding-agent benchmarks while claiming a lower cost per fix
XFreeze · x · 2026-07-26
Grok 4.5 is being pitched as a new efficiency leader in coding and agentic work, with benchmark screenshots showing it on top across several real-world tests.
- It ranks #1 on Long-Horizon Terminal-Bench for both mean reward and strict binary pass rate.
- It also leads the $5 workload category and scores 91.3% on VulcanBench v3, completing 21 of 23 tasks.
- On SignalDesk V1, it reaches 99% by fixing 69 of 70 tasks at about $0.074 per fix.
- The post claims Grok 4.5 is now seeing the biggest week-over-week usage jump inside Augment Code’s Cosmos picker.
The framing is that Grok 4.5 is not just competitive on capability, but unusually strong on cost efficiency and real-world developer adoption.
More from coding & agent
- Coding agents could speed up game-economy tuning as software value turns more winner-take-most — trq212 · 2026-07-26
- CodeInspectus adds local AI-code security scanning to Codex through MCP — hibzy7 · 2026-07-26
- Open Minis open-sources its full iOS and Android on-device agent stack — dotey · 2026-07-26
- Sam Altman says Codex was a “kamikaze mission” to catch Claude Code — soumitrashukla9 · 2026-07-26
- A chat orchestrator now drives Claude Code, Codex and Cursor agents in one place — kelvinbksoh · 2026-07-26
- Fitter turns scraping into declarative MCP configs and ships a browser playground — PyxRu · 2026-07-26