Grok 4.5 Takes Second Place on SWE Leaderboard
elonmusk · x · 2026-07-11
Grok 4.5 scored Pass@1 51.2% ±6.0 on the real-world software engineering benchmark APEX-SWE, ranking second behind Fable 5's 65.5% ±6.2.
In specific subcategories, it ranked first in Integration tasks with 65.0% Pass@1 and second in Observability with 37.3% Pass@1. The post also notes that these multi-step build and diagnostic/debugging tasks perfectly align with the agentic coding workflows Grok 4.5 targets. Compared to Grok 4's 21.0%, this represents a 30.2 percentage point improvement within a year.
Related event: Grok 4.5 Ranks No. 2 on APEX-SWE(2 posts)→
More from coding & agent
- A GLP1R variant may explain stronger Ozempic weight loss, and the team built an agent workflow — julia_kiseleva · 2026-07-21
- A Claude-coded Chrome extension shames you with a private jet when you open YouTube — alex_verem · 2026-07-21
- A curated TTS list for voice agents tracks latency, cancellation, and evals — mahimairaja · 2026-07-21
- Harness engineering is emerging as the execution layer for reliable AI agents — Pavan_Belagatti · 2026-07-21
- DevFest Lisbon keynote will cover Google AI Studio’s latest vibe coding and agentic AI features — gerardsans · 2026-07-21
- Daniel Hanchen’s 2-hour workshop covers open models, reward hacking and RL — danielhanchen · 2026-07-21