Grok 4.7 jumps on coding index but burns 81k tokens per task, 2x its predecessor

ArtificialAnlys · x · 2026-09-22

Artificial Analysis' evaluation shows Grok 4.7 (xhigh) scoring 46 on the Intelligence Index while consuming 81k output tokens per task — more than double Grok 4.6 (xhigh)'s 36k, versus 60k for Muse Spark 1.3 (max) and 27k for GPT-6 Astra (max).

Using Grok Build, its first-party coding agent, the model scores 56 on the Coding Agent Index (up from 47), improving across all components: DeepSWE v1.1 65%→73%, Terminal-Bench 4.0 18%→33%, SWE-Atlas-QnA 58%→63%. Intelligence gains come with significantly higher token usage.

Related event: Grok 4.7 Scores 46 on Intelligence Index, Enters Top Four Labs at Double Token Cost(7 posts)→

Original post →

More from Models

Models channel →