Grok 4.5 tops VulcanBench v3 with 91.3% on real repo engineering tasks
XFreeze · x · 2026-07-23
Grok 4.5 tops VulcanBench v3 for repository-level software engineering
- A screenshot of BenchLM’s public scoreboard shows Grok 4.5 at 91.3%, ahead of GPT-5.6 Sol at 87.0% and Claude Fable 5 at 87.0%.
- VulcanBench v3 is described as a demanding benchmark for real-world repository-level engineering tasks, not simple code snippets.
- The tasks are based on real merged pull requests across Python, Rust, TypeScript, JavaScript, and Go.
- Solutions are checked with deterministic hidden tests in isolated environments.
- The post says Grok 4.5 solved 21 of 23 tasks, claiming it can handle multi-file changes, debugging, and harder engineering work that needs more than one-shot answers.
Related event: Grok 4.5 Tops VulcanBench Software Engineering Leaderboard(2 posts)→
More from coding & agent
- Mapping Open-Source AI with Codex Agents: Mask2Former Found Lagging Behind SOTA — NielsRogge · 2026-07-23
- New tool manages fleets of agents across Claude Code, Codex and Cursor — dee_hw · 2026-07-23
- Dan Vega is prototyping Figma thumbnail concepts with Claude Code Skills — therealdanvega · 2026-07-23
- Databricks explains when to use Genie Agents, Knowledge Assistant, and a Supervisor — CautiousUse8597 · 2026-07-23
- Sentence Transformers 5.6.1 fixes a silent flash-attention regression in RoBERTa models — tomaarsen · 2026-07-23
- sentence-transformers 5.6.1 fixes Flash Attention embedding bug in XLM-R models — tomaarsen · 2026-07-23