Grok 4.5 tops several coding-agent benchmarks while claiming a lower cost per fix
XFreeze · x · 2026-07-26
Grok 4.5 is being pitched as a new efficiency leader in coding and agentic work, with benchmark screenshots showing it on top across several real-world tests.
- It ranks #1 on Long-Horizon Terminal-Bench for both mean reward and strict binary pass rate.
- It also leads the $5 workload category and scores 91.3% on VulcanBench v3, completing 21 of 23 tasks.
- On SignalDesk V1, it reaches 99% by fixing 69 of 70 tasks at about $0.074 per fix.
- The post claims Grok 4.5 is now seeing the biggest week-over-week usage jump inside Augment Code’s Cosmos picker.
The framing is that Grok 4.5 is not just competitive on capability, but unusually strong on cost efficiency and real-world developer adoption.
More from coding & agent
- Inspired by OpenAI's 10,000-agent run, dev open-sources a crowdsourced agent problem-solving platform — Benjaminsen · 2026-09-11
- Lucid: open-source Mac app keeps your laptop awake only while AI agents run — Pitiful_Hedgehog_600 · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11