Grok 4.5 Tops FrontierSWE Leaderboard
aman_madaan · x · 2026-07-16
Cited benchmark results indicate that Grok 4.5 ranked 2nd overall in implementation and performance on Proximal's FrontierSWE evaluation, and 1st in research capabilities.
The referenced leaderboard notes that it places just behind Claude Fable 5, but ahead of Opus 4.8, GPT-5.5, and GLM-5.2. The core takeaway here is a specific model leaderboard update rather than a discussion on products or workflows.
Related event: Grok 4.5 Rises to No. 2 on FrontierSWE(5 posts)→
More from Models
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11