Grok 4.6 Benchmarks: Leads in Coding but Trails in SWE Tasks

kimmonismus · x · 2026-08-12

Users have praised Grok 4.6 for showing an absolutely insane jump in performance.

According to xAI's benchmarks, Grok 4.6 matches GPT-5.6 Sol on the AA Intelligence Index at 61 and leads it on CursorBench, FrontierCode, and AA-Briefcase.

However, it still trails GPT-5.6 Sol on DeepSWE and Terminal-Bench. xAI noted that the model received a longer supplemental training run, regenerated SFT trajectories, and implemented agentic reinforcement learning.

Related event: xAI Launches Grok 4.6: Top Benchmark Performance at Low Cost(27 posts)→

Original post →

More from Models

Models channel →