Grok 4.5 Tops FrontierSWE Leaderboard

aman_madaan · x · 2026-07-16

Cited benchmark results indicate that Grok 4.5 ranked 2nd overall in implementation and performance on Proximal's FrontierSWE evaluation, and 1st in research capabilities.

The referenced leaderboard notes that it places just behind Claude Fable 5, but ahead of Opus 4.8, GPT-5.5, and GLM-5.2. The core takeaway here is a specific model leaderboard update rather than a discussion on products or workflows.

Related event: Grok 4.5 Rises to No. 2 on FrontierSWE(5 posts)→

Original post →

More from Models

Models channel →