AutomationBench: GPT-5.5 vs Claude Opus 4.8 Work Styles

ArtificialAnlys · x · 2026-07-07

Models exhibit significantly different work styles on AutomationBench. GPT-5.5(xhigh) leans towards dense operations, averaging 49 tool calls across 25 rounds per task. Claude Opus 4.8(max) is more deliberate, concentrating 35 tool calls within 14 rounds and recording fewer guardrail violations (0.55 vs. 0.66 per task).

Grok 4.3(high) uses the fewest rounds (13), but underperforms compared to models that persist to the end, as it tends to declare tasks complete prematurely rather than actually finishing them efficiently.

Related event: AutomationBench-AA Launches to Evaluate AI Agents on SaaS Workflows(7 posts)→

Original post →

More from Models

Models channel →