AutomationBench-AA: new benchmark tests agents on 657 real-world SaaS workflows
gordic_aleksa · x · 2026-09-18
Artificial Analysis launched AutomationBench-AA, measuring agentic task completion across 657 tasks in six business domains (Finance, HR, Marketing, Operations, Sales, Support) within simulated environments for Gmail, Sheets, Slack, Salesforce, Zendesk, Jira, and HubSpot.
- Unlike Zapier's own leaderboard, the headline metric is the average share of task objectives completed without guardrail violations
- Tasks combine cross-app coordination, autonomous REST API discovery, and policy adherence
- Built on real workflow patterns from Zapier's platform; paper available on arXiv
More from coding & agent
- 5 Code Review Skills Tested Across 30 Sessions: Vercel's Wins Overall — bootstrapper-919 · 2026-09-18
- Claude Code 2.1.275 ships 96 CLI changes, adds send-now shortcut and secret stripping — ClaudeCodeLog · 2026-09-18
- Dosu Cut Agent Debugging Time 90% Using Logfire MCP Across 697k+ Runs — samuelcolvin · 2026-09-18
- Pydantic Logfire's Sleeper Feature: Every Trace Is Just a Row in a Postgres Table — samuelcolvin · 2026-09-18
- Underrated MCP Use: Giving Coding Agents Like Claude Code Real Production Telemetry — samuelcolvin · 2026-09-18
- Jev Ditches Autoregression: A Model That Only Outputs Structured Decisions — karminski3 · 2026-09-18