Grok 4.5 Rises to No. 2 on FrontierSWE
Elon Musk said on July 15 that Grok 4.5 had climbed to second place on the FrontierSWE software engineering benchmark, making benchmark ranking the center of this discussion cluster. The result drew attention because several posters took it as a sign that Grok 4.5 has entered the current top tier of software engineering models.
Key details
According to Musk, Grok 4.5 ranked No. 2 on FrontierSWE. A relayed post shared by @amanmadaan added more detail from Proximal’s FrontierSWE evaluation: Grok 4.5 was No. 2 on the combined implementation-and-performance dimension, behind Claude Fable 5, and No. 1 on the research dimension.
Reactions and interpretations
@XFreeze argued that Grok 4.5’s coding ability is now in the first tier and said its benchmark standing puts it ahead of Claude Opus 4.8 and GPT-5.5 in this comparison. He also highlighted speed and token efficiency as notable strengths. In a separate post, he further stressed that Grok 4.5 achieved a top result while using significantly fewer tokens than Fable 5, framing efficiency as one of the model’s main selling points in this round of discussion.
2026-07-15 ~ 2026-07-16 · 5 related posts
- Episode 1: Musk Announces Grok 4.5 with 1.5T Parameters and Enhanced Coding(2026-07-07, 25 posts)
- Episode 2: Prediction Markets Strongly Price In Grok 4.4 Release(2026-07-07, 2 posts)
- Episode 3: Grok 4.5 Released with Focus on Coding and Low Cost(2026-07-08, 61 posts)
- Episode 4: xAI Launches Grok 4.5: Coding and Agent Focus to Rival Opus(2026-07-09, 55 posts)
- Episode 5: Grok 4.5 Receives Widespread Praise for Speed and Coding(2026-07-09, 13 posts)
- Episode 6: Grok 4.5 Released, Ranks 6th on Vals Index(2026-07-09, 2 posts)
- Episode 7: Grok 4.5 Integrates into Notion(2026-07-09, 2 posts)
- Episode 8: Musk Highlights Grok 4.5's High Performance and Low Cost(2026-07-09, 3 posts)
- Episode 9: Grok 4.5 Tops Multiple Professional AI Benchmarks(2026-07-10, 5 posts)
- Episode 10: Grok 4.5 Tops Coding Benchmark with High Token Efficiency(2026-07-10, 3 posts)
- Episode 11: Grok 4.5 Opens Free Tier and Integrates with Developer Tools(2026-07-10, 9 posts)
- Episode 12: Grok 4.5 Joins Perplexity as Orchestrator, Tops WANDR Benchmark(2026-07-11, 10 posts)
- Episode 13: Grok 4.5 Shows Significant Improvements in Coding Capabilities(2026-07-11, 2 posts)
- Episode 14: Grok 4.5 Evaluations: Strong Cost-Performance in Mid-to-High Budgets(2026-07-11, 5 posts)
- Episode 15: Grok 4.5 Gains Praise for Speed and Top-Tier Performance(2026-07-11, 2 posts)
- Episode 16: Grok 4.5 First Tests: Faster, Cheaper, and Entering Coding Workflows(2026-07-12, 8 posts)
- Episode 17: Musk Says Grok 4.5 Beats Fable on Some Coding Benchmarks(2026-07-13, 3 posts)
- Episode 18: Grok 4.5 Tops Coding Q&A Benchmark with New Collaborative Release(2026-07-13, 3 posts)
- Episode 19: Grok 4.5 Draws Praise for Speed Lead(2026-07-14, 2 posts)
- Episode 20: Grok 4.5 shifts focus to coding and agent workflows(2026-07-15, 12 posts)
Primary sources
- Musk: Grok 4.5 Ranks Second on FrontierSWE Benchmark — elonmusk ·
- Grok 4.5 Hits No.2 on Coding Leaderboard — XFreeze · 2026-07-15
- Grok 4.5 Shows Strong Coding Efficiency — XFreeze · 2026-07-15
- Grok 4.5 Climbs to #2 on FrontierSWE — elonmusk · 2026-07-15
- [source] Musk: Grok 4.5 Ranks Second on FrontierSWE Benchmark — elonmusk · 2026-07-15
- Grok 4.5 Tops FrontierSWE Leaderboard — aman_madaan · 2026-07-16