Grok 4.5 Tops VulcanBench, Beating Claude and GPT in Coding Accuracy and Cost

XFreeze · x · 2026-08-05

Grok 4.5 ranked #1 on the VulcanBench benchmark, solving 21 out of 23 (91.3%) real-world merged PR tasks across Python, Rust, TypeScript, JavaScript, and Go.

Beyond accuracy, Grok 4.5 demonstrated superior economics at just $0.32 per solved task. Competitors like Claude Fable 5 and GPT-5.6 Sol peaked at 20/23, while Kimi K3 required a two-hour extended budget to match the score, costing $1.37 per task.

Original post →

More from Models

Models channel →