AutomationBench-AA Launches to Evaluate AI Agents on SaaS Workflows
Artificial Analysis, in collaboration with Zapier, has launched AutomationBench-AA, a new benchmark designed to evaluate how well AI agents automate real-world SaaS workflows. The release is accompanied by an arXiv paper, a GitHub repository, and a Zapier leaderboard, establishing a standardized test for model automation and tool-calling capabilities.
Evaluation Mechanics and Task Difficulty
AutomationBench-AA consists of 657 workflow automation tasks spanning six business domains: finance, HR, marketing, operations, sales, and support. Models are required to orchestrate workflows via REST API across 40 simulated SaaS application environments, with programmatic scoring based on the final system state. Task difficulty varies significantly by domain, with financial workflows proving the most resistant to automation—models collectively completed only about one-third of financial goals, roughly half the completion rate of customer support and operations (around 60%).
Model Performance and Cost Discrepancies
Models exhibit highly distinct operational styles. For instance, GPT-5.5 (xhigh) favors intensive action, averaging 49 tool calls across 25 turns per task. In contrast, Claude Opus 4.8 (max) operates more cautiously, executing 35 tool calls within just 14 turns. The cost per task also spans over an order of magnitude: DeepSeek V4, Gemini 3.1 Flash-Lite, and Qwen3.7 Plus cost less than 5 cents per task, whereas Claude Opus 4.8 (max) approaches 1.5 USD. However, higher cost does not guarantee supremacy, as leading models are not always the most expensive.
Guardrail Violations and Efficiency
Beyond simply completing objectives, agents must avoid violating guardrails that represent business rules. Evaluations show that all assessed models trigger guardrail violations upon release. When ranking efficiency adjusted for violations (valid goals completed per violation), Gemini 3.5 Flash leads with 15.0 goals per violation, followed closely by Claude Opus 4.8 (max) at 13.5. Claude demonstrated fewer overall guardrail violations per task during its operations.
2026-07-07 ~ 2026-07-07 · 7 related posts
Primary sources
- Zapier AutomationBench-AA: SaaS Workflow Automation Leaderboard for AI Agents — ArtificialAnlys ·
- Zapier AutomationBench-AA: 657 Workflow Automation Tasks — ArtificialAnlys ·
- AutomationBench-AA: A New Benchmark for Automation Released — ArtificialAnlys ·
- [source] Zapier AutomationBench-AA: SaaS Workflow Automation Leaderboard for AI Agents — ArtificialAnlys · 2026-07-07
- [source] Zapier AutomationBench-AA: 657 Workflow Automation Tasks — ArtificialAnlys · 2026-07-07
- AutomationBench: Efficiency Rankings Adjusted by Guardrail Violations — ArtificialAnlys · 2026-07-07
- AutomationBench: Financial Workflows Are the Hardest to Automate — ArtificialAnlys · 2026-07-07
- AutomationBench: Per-Task Model Costs Span an Order of Magnitude — ArtificialAnlys · 2026-07-07
- [source] AutomationBench-AA: A New Benchmark for Automation Released — ArtificialAnlys · 2026-07-07
- AutomationBench: GPT-5.5 vs Claude Opus 4.8 Work Styles — ArtificialAnlys · 2026-07-07