Grok 4.7 sweeps top 3 on VulcanBench Frontier v4 repo-level coding benchmark

XFreeze · x · 2026-10-06

Grok 4.7 took all top 3 spots on VulcanBench Frontier v4, beating Fable 5.1, Opus 5.5, GPT-6 Astra, GPT-6.1 Sol and other frontier models:

The benchmark tests repository-level software engineering rather than isolated snippets: 23 behavioral-reconstruction tasks where models rebuild legacy software to match actual behavior, graded by hidden functional tests plus security, lint/complexity and code quality checks.

Original post →

More from coding & agent

coding & agent channel →