Grok 4.5 Excels at Agentic Tasks
ArtificialAnlys · x · 2026-07-09
Artificial Analysis notes that Grok 4.5 performs exceptionally well in agent tasks such as knowledge work, terminal operations, and customer service. It matches or exceeds Claude Opus 4.8 and GPT-5.5 on GDPval-AA v2, τ³-Banking, and Terminal-Bench v2.1.
Related event: Grok 4.5 Released with Focus on Coding and Low Cost(61 posts)→
More from coding & agent
- Fireworks says Kimi K3 handles 72–96% of agent traffic at up to 50x lower cost — eliebakouch · 2026-07-22
- Poolside releases 118B open Laguna S 2.1 for single-node local deployment — IanAndrewsDC · 2026-07-22
- AutoIndex suggests AI may improve by optimizing executable programs, not just weights — mrdrozdov · 2026-07-22
- AutoIndex lifts CRUMB recall by 8.4% without changing retrievers or embeddings — mrdrozdov · 2026-07-22
- AutoIndex defines a representation program as the logic behind search indexing — mrdrozdov · 2026-07-22
- Hermes adds rollback checkpoints and a Yolo mode for destructive coding agents — alexcovo_eth · 2026-07-22