Grok 4.5 tops VulcanBench v3 with 91.3% on real repo engineering tasks
XFreeze · x · 2026-07-23
Grok 4.5 tops VulcanBench v3 for repository-level software engineering
- A screenshot of BenchLM’s public scoreboard shows Grok 4.5 at 91.3%, ahead of GPT-5.6 Sol at 87.0% and Claude Fable 5 at 87.0%.
- VulcanBench v3 is described as a demanding benchmark for real-world repository-level engineering tasks, not simple code snippets.
- The tasks are based on real merged pull requests across Python, Rust, TypeScript, JavaScript, and Go.
- Solutions are checked with deterministic hidden tests in isolated environments.
- The post says Grok 4.5 solved 21 of 23 tasks, claiming it can handle multi-file changes, debugging, and harder engineering work that needs more than one-shot answers.
Related event: Grok 4.5 Tops VulcanBench Software Engineering Leaderboard(2 posts)→
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- ARRM targets silent economic regressions in AI agents that functional tests miss — Beautiful_Belt_601 · 2026-09-11