Grok 4.7 jumps on coding index but burns 81k tokens per task, 2x its predecessor
ArtificialAnlys · x · 2026-09-22
Artificial Analysis' evaluation shows Grok 4.7 (xhigh) scoring 46 on the Intelligence Index while consuming 81k output tokens per task — more than double Grok 4.6 (xhigh)'s 36k, versus 60k for Muse Spark 1.3 (max) and 27k for GPT-6 Astra (max).
Using Grok Build, its first-party coding agent, the model scores 56 on the Coding Agent Index (up from 47), improving across all components: DeepSWE v1.1 65%→73%, Terminal-Bench 4.0 18%→33%, SWE-Atlas-QnA 58%→63%. Intelligence gains come with significantly higher token usage.
More from Models
- Grok 4.7 falls to #24 on Vals Index, down 5 points from Grok 4.6 — scaling01 · 2026-09-22
- Game Theory of Model Launch Dates: Launching Early Admits Your Model Is Weaker — cocktailpeanut · 2026-09-22
- Liquid AI's LFM2.5 tops mobile benchmarks: 2.32GB memory, 8s latency on iPhone 17 Pro — maximelabonne · 2026-09-22
- Jev reportedly does tensor logic under the hood: differentiable IF args, no wasted gen tokens — StewartalsopIII · 2026-09-22
- Multilingual Model Laya Trending on Hugging Face — convaiinnovations · 2026-09-22
- Terminal Bench 4.0: GLM-5.3, Qwen 3.8, Muse 1.3 and DeepSeek v4.1 Lead the Chart — himanshustwts · 2026-09-22