Grok 4.6 ranked #2 on AutomationBench-AA, beating Claude, GPT-5.6 and Kimi
XFreeze · x · 2026-09-08
X user XFreeze reports that Grok 4.6 ranks #2 on the updated AutomationBench-AA, outperforming Claude Fable 5.1, GPT-5.6, Kimi K3 and others. The benchmark tests real agentic workflows across SaaS tools rather than static Q&A, suggesting Grok is getting significantly stronger at actually executing work. Note: third-party report without a linked source; unverified.
More from Models
- Users With Cyber Access Keep Hitting Claude Guardrails Daily, Sparking Overblocking Complaints — IgorCarron · 2026-09-08
- Fable vs astra: One Asks Why, the Other Dives Straight Into Code — Liu_eroteme · 2026-09-08
- DeepSeek spotted gray-testing two V4-Flash-Vision models, one faster but weaker — teortaxesTex · 2026-09-08
- GPT-6 Astra demos roundup: livestreams, retro games, 3D scenes and full websites — aitrendz_xyz · 2026-09-08
- DeepSeek V4.1 Flash beta goes live with new architecture, native multimodality, up to 507 tokens/s — 智东西 · 2026-09-08
- OpenAI API users report day-long outages: searches hang 30 minutes then drop with no output — msch6873 · 2026-09-08