Grok 4.5 Takes Second Place on SWE Leaderboard

elonmusk · x · 2026-07-11

Grok 4.5 scored Pass@1 51.2% ±6.0 on the real-world software engineering benchmark APEX-SWE, ranking second behind Fable 5's 65.5% ±6.2.

In specific subcategories, it ranked first in Integration tasks with 65.0% Pass@1 and second in Observability with 37.3% Pass@1. The post also notes that these multi-step build and diagnostic/debugging tasks perfectly align with the agentic coding workflows Grok 4.5 targets. Compared to Grok 4's 21.0%, this represents a 30.2 percentage point improvement within a year.

Related event: Grok 4.5 Ranks No. 2 on APEX-SWE(2 posts)→

Original post →

More from coding & agent

coding & agent channel →