Grok 4.5 Excels at Agentic Tasks

ArtificialAnlys · x · 2026-07-09

Artificial Analysis notes that Grok 4.5 performs exceptionally well in agent tasks such as knowledge work, terminal operations, and customer service. It matches or exceeds Claude Opus 4.8 and GPT-5.5 on GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1.

Related event: Grok 4.5 Released with Focus on Coding and Low Cost(61 posts)→

Original post →

More from coding & agent

coding & agent channel →