Grok 4.5 tops VulcanBench on accuracy and cost efficiency
XFreeze · x · 2026-07-22
VulcanBench ranks Grok 4.5 first on both accuracy and cost efficiency.
- At medium effort, Grok 4.5 solved 21/23 tasks with a 91.3% pass rate at a total cost of $6.67.
- That works out to about $0.32 per solved task, versus $0.61 for the best Claude Fable 5 run and $0.80 for the best GPT-5.6 Sol run.
- The benchmark screenshot also shows Grok 4.5 holding the top spot on the efficiency frontier, with higher effort increasing cost but not improving the 21/23 result.
The post argues that Grok 4.5 is the best tradeoff on the board: more accurate than Fable 5 and GPT-5.6 Sol, while costing substantially less per solved task.
More from Models
- Google Exec Seeks Feedback on Gemini 3.6 Flash & 3.5 Flash-Lite Performance — patloeber · 2026-07-22
- Gemini 3.6 Flash is 2x faster and 18% cheaper, but independent tests say it is not smarter — etherd0t · 2026-07-22
- Critic says OpenAI incident coverage confuses bad reward functions with autonomy — ambaonadventure · 2026-07-22
- Google Launches Gemini 3.5 Flash Cyber Model for Security Teams — pushmeet · 2026-07-22
- Rumor says GPT-5.6 Sol could hit 750 tok/s after Cerebras upgrades — haider1 · 2026-07-22
- LeCun reposts Hugging Face’s case for open-weight models in cyber defense — ylecun · 2026-07-22