Grok 4.7 review: top-four capability, doubled token use sparks cost debate
Grok 4.7 underwent concentrated third-party evaluations on 09-22. Artificial Analysis gave it an Intelligence Index score of 46 (xhigh tier), 2 points above Grok 4.6, marking the first time xAI cracked the top four labs on the index. The overall verdict: capability rose significantly but token consumption ballooned, so real-world value depends on use case—sparking community debate.
Confirmed
- Artificial Analysis evaluation: Intelligence Index 46 (vs. 44 for Grok 4.6); hallucination rate down from 34% to 29%; output speed around 188 tokens/second; average Intelligence Index task takes about 7.1 minutes
- AA-Briefcase (multi-hour office-task benchmark): 1657 Elo, up +111 over Grok 4.6 (high), second only to the Claude family; analysis quality Elo rose to 1994, leading GPT-6 Astra by +88 Elo; leads GPT-6 Astra by +153 Elo on the GDPval-AA professional-deliverables benchmark
- On AA-Briefcase, Grok 4.7 xHigh ranks second at 58%, just 1 point behind leader Claude Fable 5.1 Max (59%), ahead of GPT-6 Astra
- Coding: CursorBench 4.0 at 46.3%, nearly matching Opus 5 Max's 46.6%; CursorBench up from 40.4% to 46.3%, EEBench from 53% to 64%; paired with the Grok Build coding agent, the Coding Agent Index rose from 47 to 56
- Cost and consumption: Grok 4.7 (xhigh) consumes roughly 81k output tokens per task, more than double Grok 4.6 (xhigh); per-task cost exceeds Astra
- Case demos: In an AA-Briefcase due-diligence scenario, Grok 4.7 independently ran through a step-by-step valuation chain and challenged a counterparty's quick estimate; in another case it analyzed three full years of financials and saw through FX-masked growth, where 4.6 used only the latest fiscal year
Unconfirmed
- User XFreeze's comparison data claims that on the same coding task, Grok 4.7 xHigh's output is roughly on par with Opus 5 Max at under half the cost, with fewer tokens and steps—this sits in tension with Artificial Analysis's cost data and reflects personal testing methodology
- ChrisGPT claims Grok 4.8 will launch before the end of November and that Grok 4.7's per-task cost exceeds GPT—these are leak-style claims
Why it matters
- Users including Angaisb and Theo (retweeting) pointed out that Grok 4.7's per-task token consumption approaches 3x Astra's and about 2.25x Grok 4.6's; under "same price, same speed" marketing, actual usage costs may be higher, and OpenAI is viewed by them as the only lab seriously caring about token efficiency
- Some users, after comparing, questioned: Grok 4.7 xhigh burns roughly 2.5x the tokens on the artificial arena for only +2 points over its predecessor—suggesting the "quality" of index gains should be measured by cost-effectiveness
- For heavy API users, xAI's new version matches frontier-model capability while unit costs rise, so model selection requires recalculating usage against their own tasks; meanwhile, its improved analytical depth on multi-hour agent tasks (e.g., the three-year financials and valuation chain cases) is a clear capability gain for professional workflows
2026-09-22 ~ 2026-09-22 · 24 related posts
- Episode 1: Grok 4.7 briefly surfaces on OpenCode Zen, launch rumored imminent(2026-09-21, 5 posts)
- Episode 2: xAI Launches Grok 4.7 with Major Coding Benchmarks Gains at Same Price(2026-09-21, 56 posts)
- Episode 3: Grok 4.7 Day One: Developers Praise Coding and Generation Gains in Hands-On Tests(2026-09-22, 15 posts)
- Episode 4: Grok 4.7 review: top-four capability, doubled token use sparks cost debate(2026-09-22, 24 posts)
- Episode 5: Grok 4.7 Reportedly Closes Gap With Top Frontier Models at Lower Cost(2026-09-22, 2 posts)
- Episode 6: Open-Source Grok 4.7 Editor Extension Passes 140K Installs(2026-09-22, 3 posts)
- Episode 7: Bug Hunt Bench: GPT-6 Astra Scores 45, Grok 4.7 Fails to Beat Grok 4.6(2026-09-22, 6 posts)
- Episode 8: Grok 4.7 vs Xiaomi MiMo-V2.6: Similar Smarts, 16x Price Gap(2026-09-22, 4 posts)
Primary sources
- Grok 4.7 scores 46 on AA Intelligence Index, enters top 4 labs but burns 81k tokens per task — ArtificialAnlys ·
- Grok 4.7's 81k output tokens per task more than double Grok 4.6's 36k — ArtificialAnlys ·
- Grok 4.7 ranks just behind Anthropic's Opus 5 on AA-Briefcase at ~50% of the cost per task — ArtificialAnlys ·
- Grok 4.7 Reportedly Uses Nearly 3x More Tokens per Task Than Astra — Angaisb_ · 2026-09-22
- [source] Grok 4.7 scores 46 on AA Intelligence Index, enters top 4 labs but burns 81k tokens per task — ArtificialAnlys · 2026-09-22
- Grok 4.7's AA-Briefcase analytical quality Elo jumps to 1994 from 1690 — ArtificialAnlys · 2026-09-22
- Grok 4.7 + Grok Build jumps to 56 on AA Coding Agent Index, up from 47 — ArtificialAnlys · 2026-09-22
- Grok 4.7 jumps on coding index but burns 81k tokens per task, 2x its predecessor — ArtificialAnlys · 2026-09-22
- [source] Grok 4.7's 81k output tokens per task more than double Grok 4.6's 36k — ArtificialAnlys · 2026-09-22
- Grok 4.7 measured at ~188 tokens/second, ~7.1 minutes per Intelligence Index task — ArtificialAnlys · 2026-09-22
- Same-prompt test: Grok 4.7 takes 32 min at $8.14 while free SWE 2 finishes in 15 min — iamfakhrealam · 2026-09-22
- Grok 4.7 cuts hallucination rate to 29% from 34% on AA-Omniscience — ArtificialAnlys · 2026-09-22
- Grok 4.7 matches Opus 5 Max on coding at less than half the cost — XFreeze · 2026-09-22
- Grok 4.7 xHigh hits 46.3% on CursorBench 4.0, matching Opus 5 Max at half the cost — XFreeze · 2026-09-22
- Grok 4.7 xHigh hits 58% on AA-Briefcase, just 1 point behind Claude Fable 5.1 Max — XFreeze · 2026-09-22
- Unverified: Grok 4.7 out now, costs more per task than GPT, says leaker — ChrisGPT · 2026-09-22
- Grok 4.7 Slammed in Benchmarks: 2.5x Tokens for Just 2 Points More — ivan_bezdomny · 2026-09-22
- [source] Grok 4.7 ranks just behind Anthropic's Opus 5 on AA-Briefcase at ~50% of the cost per task — ArtificialAnlys · 2026-09-22
- AA example: Grok 4.7 independently runs valuation chain and flags divergence from deal partner — ArtificialAnlys · 2026-09-22
- Grok 4.7 example: three-year financial analysis exposes currency-masked growth stall — ArtificialAnlys · 2026-09-22
- Grok 4.7 jumps to 46.3% on CursorBench and 64% on EEBench, keeping the same $2/$6 per million token pricing — FinanceYF5 · 2026-09-22
- Grok 4.7 looks pricier than before, now costlier than Astra on Artificial Analysis — steipete · 2026-09-22
- Grok 4.7 Token Efficiency Falls 30-80% Short of Claims, Real-World Costs 2x Grok 4.6 — flowersslop · 2026-09-22
- Grok 4.7 beats GPT-6 Astra on professional work, leading by 153 Elo on GDPval-AA — XFreeze · 2026-09-22
- Grok 4.7 costs $2.73 per task vs $1.86 for a +2 point AA index gain — koltregaskes · 2026-09-22
- Grok 4.7 lands: near Opus 5 on AA-Briefcase at ~50% cost per task — NicoVerderosa · 2026-09-22
1 near-duplicate retellings: Angaisb_