Grok 4.7 sweeps top 3 on VulcanBench Frontier v4 repo-level coding benchmark
XFreeze · x · 2026-10-06
Grok 4.7 took all top 3 spots on VulcanBench Frontier v4, beating Fable 5.1, Opus 5.5, GPT-6 Astra, GPT-6.1 Sol and other frontier models:
- xHigh: 93.15, passing 23/23 tasks
- High: 92.71, 22/23 passed
- Medium: 92.30, 21/23 passed
The benchmark tests repository-level software engineering rather than isolated snippets: 23 behavioral-reconstruction tasks where models rebuild legacy software to match actual behavior, graded by hidden functional tests plus security, lint/complexity and code quality checks.
More from coding & agent
- Most Lawyers Use AI Legal Tools Only at Basic One-Shot Prompt Level, Says Attorney — jkubicki · 2026-10-06
- GPT-2 × Codex builds 'alien eye' tracker web app that admits when it can't see — MikePFrank · 2026-10-06
- Codex tasks widget broken? Editing it to select a specific host fixes it for now — Dimillian · 2026-10-06
- xAI TypeScript SDK hits v0.2.2: retryBeforeOutput now retries create() without stream — tetsuoai · 2026-10-06
- Hugging Face Kernels quickstart: load GPU-optimized kernels in one line — ariG23498 · 2026-10-06
- Lovable CEO Anton Osika teases building an entire company with AI in 2026 — santiviquez · 2026-10-06