Grok 4.5 Takes Second Place on SWE Leaderboard
elonmusk · x · 2026-07-11
Grok 4.5 scored Pass@1 51.2% ±6.0 on the real-world software engineering benchmark APEX-SWE, ranking second behind Fable 5's 65.5% ±6.2.
In specific subcategories, it ranked first in Integration tasks with 65.0% Pass@1 and second in Observability with 37.3% Pass@1. The post also notes that these multi-step build and diagnostic/debugging tasks perfectly align with the agentic coding workflows Grok 4.5 targets. Compared to Grok 4's 21.0%, this represents a 30.2 percentage point improvement within a year.
Related event: Grok 4.5 Ranks No. 2 on APEX-SWE(2 posts)→
More from coding & agent
- SymKit MCP: 44 tools for AI agents to verify symbolic derivations — Foreign-Specific-604 · 2026-09-11
- GitHub Copilot team routes user bug reports to an AI agent via Slack — marlene_zw · 2026-09-11
- Scanning 23 agent sessions, a dev found 3 silent failure modes in memory systems — No_Advertising2536 · 2026-09-11
- eslint-plugin-react v8.0.2 adds 4 checks for React 19.3, supports ESLint 10 and Biome — viglovikov · 2026-09-11
- Arkon: open-source MCP server turns enterprise SOPs into a traceable LLM knowledge wiki — tom_doerr · 2026-09-11
- Cheaper OpenAI Agents API alternative: sandbox service undercutting E2B by 46% — airesearch12 · 2026-09-11