Grok 4.6 Tops Knowledge-Work, Productivity, and Legal Benchmarks
elonmusk · x · 2026-08-13
Elon Musk retweeted news regarding Grok 4.6's latest benchmark performance. The model reportedly leads on the two strongest knowledge-work and real-world productivity benchmarks (GDPVal-AA and AA-Briefcase), as well as on legal benchmarks, while remaining highly competitive in coding-agent and general intelligence metrics.
More from Models
- Dev Reflects on AI Coding: Models Make Basic Reasoning Errors, Hand-Coding Wins — jsuarez · 2026-08-13
- Grok Hits SOTA on Databricks' OfficeQA Pro V2 Evaluation — GavinSBaker · 2026-08-13
- Grok 4.6 hits Vercel AI Gateway with 500K token context window — soleio · 2026-08-13
- xAI Grants SuperGrok Users a Free Weekly Usage Reset — XFreeze · 2026-08-13
- Krea, FLUX, and MiniMax Open-Weights: Promises vs Reality — felixsanz · 2026-08-13
- Anthropic Slammed by Users for Extra 'Fast Mode' Fees on Claude Code — Neel_MynO · 2026-08-13