Grok 4.5 Tops VulcanBench, Beating Claude and GPT in Coding Accuracy and Cost
XFreeze · x · 2026-08-05
Grok 4.5 ranked #1 on the VulcanBench benchmark, solving 21 out of 23 (91.3%) real-world merged PR tasks across Python, Rust, TypeScript, JavaScript, and Go.
Beyond accuracy, Grok 4.5 demonstrated superior economics at just $0.32 per solved task. Competitors like Claude Fable 5 and GPT-5.6 Sol peaked at 20/23, while Kimi K3 required a two-hour extended budget to match the score, costing $1.37 per task.
More from Models
- OpenAI's gpt-5.6-luna So Cost-Effective It Constantly Overloads Servers — _lewtun · 2026-08-05
- Opus 5 Language Degradation: Anthropic's Model Accused of Being a 'Jargon Douche' — TheTuringPost · 2026-08-05
- NVIDIA Releases Alpamayo 2 Super: A 34B Parameter VLA Model for Autonomous Driving — drmapavone · 2026-08-05
- AI Needs a Nap? Developer Exposes Hilarious GPT Thinking Block Hallucination — cantrell · 2026-08-05
- engyai Launches Cheapest Kimi K3 API on OpenRouter, Cutting Costs by 50% — const_reborn · 2026-08-05
- LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context — helloiamleonie · 2026-08-05