Grok 4.5 Tops Automation Benchmark

ArtificialAnlys · x · 2026-07-09

The AutomationBench-AA benchmark shows Grok 4.5 taking first place in automated workflow tasks with a score of 51%, slightly ahead of Claude Fable 5 at 49% and Claude Opus 4.8 at 48%. The post also notes that its cost per task is about a quarter of its competitors, and it is the first model to complete over half of the workflow goals without violating business rules. Maintained independently by AutomationBench-AA, the benchmark tests whether AI agents can automate 657 tasks across 40 simulated SaaS environments like Gmail and Google Sheets.

Related event: Grok 4.5 Tops Automation Benchmark(2 posts)→

Original post →

More from coding & agent

coding & agent channel →