Grok 4.5 Leads Non-Anthropic Models in Evaluation
ArtificialAnlys · x · 2026-07-10
Grok 4.5 was rated the best-performing non-Anthropic model on AA-Briefcase, scoring 1328, a 578-point increase over Grok 4.3.
The evaluation also highlights its cost and speed advantages, averaging $1.12 and 12.4 minutes per task, respectively outperforming multiple compared models.
Related event: Grok 4.5 Released with Focus on Coding and Low Cost(61 posts)→
More from Models
- Flash 3.8 Review: Great When Working, but Stuck in Silent Token-Burning Loops — brandon_galang · 2026-09-03
- Claude Max 20x buyer says weekly limits, not the 5-hour window, are the real bottleneck; r/ClaudeAI deleted his post — conorearly · 2026-09-03
- Users say Gemini 3.8 Flash regressed for agentic tasks, calling 3.7 Flash calmer and faster — Scobleizer · 2026-09-03
- Scale CEO Praises Muse Spark 1.3's Interactive 3D Japanese Garden Scene as Huge Leap — alexandr_wang · 2026-09-03
- AA Predictions List Sparks Buzz With Mysterious 'DeepSeekV5-Preview: 58' Entry — teortaxesTex · 2026-09-03
- Gemini 3.8 Flash praised for quality but keeps hitting infinite loops in Cursor — brandon_galang · 2026-09-03