Grok 4.6 Debuts Strong on AA-Briefcase, Trailing Only Claude Opus 5
ArtificialAnlys · x · 2026-08-12
On Artificial Analysis's private benchmark for long-horizon agentic knowledge work tasks (AA-Briefcase), xAI's Grok 4.6 debuted with an Elo of 1577, finishing just behind Claude Opus 5.
The results indicate that Grok 4.6 performs consistently and strongly across multiple dimensions, including rubric grading, presentation, and analytical quality.
More from Models
- Grok 4.6 Pricing Becomes Highly Competitive, Pressuring Cursor Subscriptions — Rasmic · 2026-08-13
- Grok 4.6 Matches Competitor Performance at 85% Discount; Next-Gen to Include Cursor Data — altryne · 2026-08-13
- Vercel AI Gateway Adds Grok and DeepSeek Models with Zero Markup — brandon_galang · 2026-08-13
- DeepSeek V4 Pro Now Available on OpenCode Go with Leaked Benchmarks — op7418 · 2026-08-13
- Grok 4.6 Tops Artificial Analysis Agentic Index, Tying Claude Opus 5 Max — XFreeze · 2026-08-13
- DeepSeek 4 Flash Price Wars Thrive as Community Awaits 4 Pro Pricing — AccBalanced · 2026-08-13