Grok 4.7 scores 46 on AA Intelligence Index, enters top 4 labs but burns 81k tokens per task
ArtificialAnlys · x · 2026-09-22
Artificial Analysis evaluated Grok 4.7, bringing xAI into the top 4 labs:
- Intelligence Index: 46, +2 over Grok 4.6 (evaluated at xhigh reasoning effort)
- Agentic knowledge work: 1657 Elo on AA-Briefcase (+111), joining Claude Opus 5 and Claude Fable 5.1 at the frontier; 1695 Elo on GDPval-AA
- Coding agents: 56 on the Coding Agent Index with Grok Build (+9), ranking 4th natively, overtaking GPT-5.6 Sol
- Token cost: 81k output tokens per task, 125% more than Grok 4.6 (36k) and 196% more than GPT-6 Astra (27k)
- Hallucination rate improved 34%→29%; Terminal-Bench 4.0 +4.5pp, AA-LCR -3.7pp
- 500k context unchanged; pricing $2/$6 per 1M tokens ($0.50 cache hits); 188 tokens/sec output
More from Models
- Reliquary-4B: A 4B math & code model trained via decentralized RL with community rollouts — const_reborn · 2026-09-22
- Users say they can't trick Jev into hallucinating — BLUECOW009 · 2026-09-22
- Measured trade-offs of three REAP-pruned Qwen3.8-Flash-Next MLX builds on Apple Silicon — MensaProdigy · 2026-09-22
- Dev claims further-optimized DeepSeek V4 NVFP4 uses 190GB of 192GB VRAM — HankYeomans · 2026-09-22
- OpenAI researcher Will Depue on why voice models still lack true realtime chat — willdepue · 2026-09-22
- Hands-on: Ling-3.0-flash-VL generates full web page code from a screenshot in 17 seconds — _jaydeepkarale · 2026-09-22