Grok 4.5 Tops Automation Benchmark
ArtificialAnlys · x · 2026-07-09
The AutomationBench-AA benchmark shows Grok 4.5 taking first place in automated workflow tasks with a score of 51%, slightly ahead of Claude Fable 5 at 49% and Claude Opus 4.8 at 48%. The post also notes that its cost per task is about a quarter of its competitors, and it is the first model to complete over half of the workflow goals without violating business rules. Maintained independently by AutomationBench-AA, the benchmark tests whether AI agents can automate 657 tasks across 40 simulated SaaS environments like Gmail and Google Sheets.
Related event: Grok 4.5 Tops Automation Benchmark(2 posts)→
More from coding & agent
- Kimi staff member builds a VR companion with Kimi Code K3 demo — dejavucoder · 2026-07-21
- OCR repo adds a JSON model directory to help agents pick the right model — strickvl · 2026-07-21
- Cross-agent system logs are dominated by questions and code proposals — nptacek · 2026-07-21
- Coding agents need better rules for when to read search summaries or full pages — RhubarbLarge2747 · 2026-07-21
- Notch says he may try vibe coding after struggling to hire good programmers — max_paperclips · 2026-07-21
- Seedance 2.0 keeps character consistency across 15+ shots with just 3 prompts — techhalla · 2026-07-21