Grok 4.6 Debuts Strong on AA-Briefcase, Trailing Only Claude Opus 5

ArtificialAnlys · x · 2026-08-12

On Artificial Analysis's private benchmark for long-horizon agentic knowledge work tasks (AA-Briefcase), xAI's Grok 4.6 debuted with an Elo of 1577, finishing just behind Claude Opus 5.

The results indicate that Grok 4.6 performs consistently and strongly across multiple dimensions, including rubric grading, presentation, and analytical quality.

Original post →

More from Models

Models channel →