AutomationBench: Efficiency Rankings Adjusted by Guardrail Violations

ArtificialAnlys · x · 2026-07-07

While completing goals, agents must also avoid violating guardrails that represent business rules. All evaluated models triggered guardrail violations upon release. Adjusting efficiency by violations helps distinguish the leaders: Gemini 3.5 Flash completed 15.0 goals per violation, while Claude Opus 4.8(max) achieved 13.5.

Related event: AutomationBench-AA Launches to Evaluate AI Agents on SaaS Workflows(7 posts)→

Original post →

More from Models

Models channel →