AutomationBench: GPT-5.5 vs Claude Opus 4.8 Work Styles
ArtificialAnlys · x · 2026-07-07
Models exhibit significantly different work styles on AutomationBench. GPT-5.5(xhigh) leans towards dense operations, averaging 49 tool calls across 25 rounds per task. Claude Opus 4.8(max) is more deliberate, concentrating 35 tool calls within 14 rounds and recording fewer guardrail violations (0.55 vs. 0.66 per task).
Grok 4.3(high) uses the fewest rounds (13), but underperforms compared to models that persist to the end, as it tends to declare tasks complete prematurely rather than actually finishing them efficiently.
Related event: AutomationBench-AA Launches to Evaluate AI Agents on SaaS Workflows(7 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11