Grok 4.6 Model Card Analysis: Big Internal Gains, Lags on Public SWE Evals

scaling01 · x · 2026-08-13

xAI recently released the model card for Grok 4.6, detailing its capabilities across coding, engineering, and knowledge work, alongside pre-deployment safety testing.

Analysis shows huge progress on benchmarks like DeepSearchQA and their internal KernelBench, and it achieves SOTA on inferenceEval. However, Grok 4.6 still lags behind the frontier on public software engineering evaluations such as Terminal Bench 3.0 and SWE Marathon v1.1.

Original post →

More from Models

Models channel →