Grok 4.7's AA-Briefcase analytical quality Elo jumps to 1994 from 1690
ArtificialAnlys · x · 2026-09-22
Artificial Analysis breaks down Grok 4.7's agentic knowledge work scores: 1657 Elo on AA-Briefcase (+111 over Grok 4.6 high), just behind Claude Opus 5 and Claude Fable 5.1. Gains are driven by analytical quality: 1994 Elo vs 1690, while presentation quality slipped slightly to 1499 from 1519. The benchmark also checks deliverables against task requirements. On GDPval-AA (documents, spreadsheets, slides), it scores 1695 vs 1605 for the predecessor.
More from Models
- Grok 4.7 falls to #24 on Vals Index, down 5 points from Grok 4.6 — scaling01 · 2026-09-22
- Game Theory of Model Launch Dates: Launching Early Admits Your Model Is Weaker — cocktailpeanut · 2026-09-22
- Liquid AI's LFM2.5 tops mobile benchmarks: 2.32GB memory, 8s latency on iPhone 17 Pro — maximelabonne · 2026-09-22
- Jev reportedly does tensor logic under the hood: differentiable IF args, no wasted gen tokens — StewartalsopIII · 2026-09-22
- Multilingual Model Laya Trending on Hugging Face — convaiinnovations · 2026-09-22
- Terminal Bench 4.0: GLM-5.3, Qwen 3.8, Muse 1.3 and DeepSeek v4.1 Lead the Chart — himanshustwts · 2026-09-22