Grok 4.5 Leads Non-Anthropic Models in Evaluation

ArtificialAnlys · x · 2026-07-10

Grok 4.5 was rated the best-performing non-Anthropic model on AA-Briefcase, scoring 1328, a 578-point increase over Grok 4.3.

The evaluation also highlights its cost and speed advantages, averaging $1.12 and 12.4 minutes per task, respectively outperforming multiple compared models.

Related event: Grok 4.5 Released with Focus on Coding and Low Cost(61 posts)→

Original post →

More from Models

Models channel →