Grok 4.6 ranked #2 on AutomationBench-AA, beating Claude, GPT-5.6 and Kimi

XFreeze · x · 2026-09-08

X user XFreeze reports that Grok 4.6 ranks #2 on the updated AutomationBench-AA, outperforming Claude Fable 5.1, GPT-5.6, Kimi K3 and others. The benchmark tests real agentic workflows across SaaS tools rather than static Q&A, suggesting Grok is getting significantly stronger at actually executing work. Note: third-party report without a linked source; unverified.

Original post →

More from Models

Models channel →