Grok 4.7 Scores 46 on Intelligence Index, Enters Top Four Labs at Double Token Cost
On September 22, Artificial Analysis published its evaluation of the Grok 4.7 (xhigh) series: it scored 46 on the Intelligence Index, 2 points above Grok 4.6, putting xAI in the top four labs on the Index for the first time. The capability gains come at a cost—token consumption has doubled, with roughly 81k output tokens per task, about twice that of rival models.
Confirmed
- Intelligence Index: Grok 4.7 (xhigh) scored 46, up 2 points from Grok 4.6, marking xAI's first entry into the top four labs on the Intelligence Index.
- Coding: Paired with xAI's own Grok Build coding agent, Grok 4.7 (xhigh) scored 56 on the Coding Agent Index, a big jump from Grok 4.6's 47, with gains across all three sub-benchmarks.
- Knowledge QA: On the AA-Omniscience sub-benchmark, the hallucination rate dropped from 34% to 29%; the full evaluation breakdown is now public.
- Agentic knowledge work: AA-Briefcase (simulating real professional tasks) scored 1657 Elo, up 111 from Grok 4.6 (high), second only to Claude-family models.
- Performance tests: Output speed was about 188 tokens/second on long prompts, with Intelligence Index tasks averaging roughly 7.1 minutes each.
- The cost: Each task consumes roughly 81k output tokens—double Grok 4.6 (xhigh) and about twice other competitors.
Why it matters
- Grok 4.7 puts xAI in the Intelligence Index top four for the first time, signaling its arrival as a first-tier lab.
- But the gains come with steep token costs; Artificial Analysis specifically flags this "pay more for capability" trade-off, an important reference point for cost-effectiveness assessments in real-world deployments.
2026-09-22 ~ 2026-09-22 · 7 related posts
Primary sources
- Grok 4.7 scores 46 on AA Intelligence Index, enters top 4 labs but burns 81k tokens per task — ArtificialAnlys ·
- Grok 4.7's 81k output tokens per task more than double Grok 4.6's 36k — ArtificialAnlys ·
- Grok 4.7 + Grok Build jumps to 56 on AA Coding Agent Index, up from 47 — ArtificialAnlys ·
- [source] Grok 4.7 scores 46 on AA Intelligence Index, enters top 4 labs but burns 81k tokens per task — ArtificialAnlys · 2026-09-22
- Grok 4.7's AA-Briefcase analytical quality Elo jumps to 1994 from 1690 — ArtificialAnlys · 2026-09-22
- [source] Grok 4.7 + Grok Build jumps to 56 on AA Coding Agent Index, up from 47 — ArtificialAnlys · 2026-09-22
- Grok 4.7 jumps on coding index but burns 81k tokens per task, 2x its predecessor — ArtificialAnlys · 2026-09-22
- [source] Grok 4.7's 81k output tokens per task more than double Grok 4.6's 36k — ArtificialAnlys · 2026-09-22
- Grok 4.7 measured at ~188 tokens/second, ~7.1 minutes per Intelligence Index task — ArtificialAnlys · 2026-09-22
- Grok 4.7 cuts hallucination rate to 29% from 34% on AA-Omniscience — ArtificialAnlys · 2026-09-22