Grok 4.5 Rises to No. 2 on FrontierSWE

Elon Musk said on July 15 that Grok 4.5 had climbed to second place on the FrontierSWE software engineering benchmark, making benchmark ranking the center of this discussion cluster. The result drew attention because several posters took it as a sign that Grok 4.5 has entered the current top tier of software engineering models.

Key details

According to Musk, Grok 4.5 ranked No. 2 on FrontierSWE. A relayed post shared by @aman_madaan added more detail from Proximal’s FrontierSWE evaluation: Grok 4.5 was No. 2 on the combined implementation-and-performance dimension, behind Claude Fable 5, and No. 1 on the research dimension.

Reactions and interpretations

@XFreeze argued that Grok 4.5’s coding ability is now in the first tier and said its benchmark standing puts it ahead of Claude Opus 4.8 and GPT-5.5 in this comparison. He also highlighted speed and token efficiency as notable strengths. In a separate post, he further stressed that Grok 4.5 achieved a top result while using significantly fewer tokens than Fable 5, framing efficiency as one of the model’s main selling points in this round of discussion.

2026-07-15 ~ 2026-07-16 · 5 related posts

Full story(20 episodes)→