Grok 4.5 reportedly solves 21 of 23 tasks on VulcanBench v3
eyishazyer · x · 2026-07-24
Coding benchmark claim
- The model reportedly solved 21 of 23 tasks on VulcanBench v3, a benchmark aimed at multi-file repository work rather than toy problems.
- It is said to reach 91.3% overall, ahead of Fable 5 and GPT-5.6 Sol.
The catch
- The thread says other benchmarks tell a different story:
- Fable 5 wins Terminal-Bench and DeepSWE.
- Opus 4.8 edges Grok on DeepSWE 1.1.
- Grok’s clear win is SWE Marathon.
- The author’s point: it is not dominating every benchmark, but it is making a strong case on one very relevant coding test.
Related event: xAI's Flagship Grok 4.5 Launches Across All Platforms Simultaneously(16 posts)→
More from Models
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07
- New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist — dair_ai · 2026-09-07
- Alexandr Wang flags Muse Spark 1.3 eval: time horizon now matches GPT-5.6 Sol and Opus 5 — alexandr_wang · 2026-09-07
- Cool presentation aside, Astra still can't nail research-level single-step reasoning — xiaosun86 · 2026-09-07
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07