Grok 4.5 Leads in Multiple Agent Tasks

ArtificialAnlys · x · 2026-07-09

Artificial Analysis states that Grok 4.5 shows strong performance in knowledge work, terminal usage, and customer service tasks. On GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1, it is on par with or surpasses Claude Opus 4.8 and GPT-5.5.

Related event: Grok 4.5 Released with Focus on Coding and Low Cost(61 posts)→

Original post →

More from coding & agent

coding & agent channel →