Artificial Analysis Intelligence Index v4.3: Terminal-Bench 4.0, Zapier's AutomationBench

SteppenAxolotl · reddit · 2026-09-08

Artificial Analysis shipped Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 (66 multi-step terminal tasks, now on mini-SWE-agent) and replacing τ³-Banking with AutomationBench-AA, a 657-task private business-workflow benchmark built on Zapier's set, lifting private-test weight to 45%. Claude Fable 5.1 and GPT-6 Astra tie at 53; GLM-5.3 and Kimi K3 lead open weights at 44; OpenAI dominates the cost-efficiency frontier.

Related event: Artificial Analysis Intelligence Index Updated to v4.3(4 posts)→

Original post →

More from Models

Models channel →