Artificial Analysis Intelligence Index v4.3: Terminal-Bench 4.0, Zapier's AutomationBench
SteppenAxolotl · reddit · 2026-09-08
Artificial Analysis shipped Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 (66 multi-step terminal tasks, now on mini-SWE-agent) and replacing τ³-Banking with AutomationBench-AA, a 657-task private business-workflow benchmark built on Zapier's set, lifting private-test weight to 45%. Claude Fable 5.1 and GPT-6 Astra tie at 53; GLM-5.3 and Kimi K3 lead open weights at 44; OpenAI dominates the cost-efficiency frontier.
Related event: Artificial Analysis Intelligence Index Updated to v4.3(4 posts)→
More from Models
- Heavy user burns 12% of ChatGPT quota in 2 hours, mulls buying a second account — BLUECOW009 · 2026-09-08
- Every tested model reinforced user delusions in mental health scenarios, with safety only in ~40% of turns — alex_verem · 2026-09-08
- LLM Museum: historical benchmarks and infographics of frontier models since GPT-4 — delusion54 · 2026-09-08
- GPT-6 Astra generates 13,038 annotations across 81 warehouse CCTV frames — luisdans · 2026-09-08
- After 3 days of testing, GPT-6 Astra looks overhyped: overfit and messy code — ivan_bezdomny · 2026-09-08
- Dev wants ChatGPT desktop but multi-account limits force him back to the TUI — kevinkern · 2026-09-08