Artificial Analysis ships Intelligence Index v4.3 with AutomationBench-AA
SonglinYang4 · x · 2026-09-08
- Artificial Analysis announced Intelligence Index v4.3
- Terminal-Bench upgraded 2.1 → 4.0; τ³-Banking replaced by AutomationBench-AA, an agentic workflow automation benchmark (based on Zapier's benchmark, with a private test set)
- Positioned as a partial rollout ahead of Index v5; each change stands on its own
- EdwardSun0909 comments: a model good across many measurements may simply be good, not benchmaxxing
Related event: Artificial Analysis Updates Intelligence Index to v4.3 with New Benchmarks(5 posts)→
More from Models
- Agent midway to AGI gives up writing full sentences, author jokes about blaming post-training — yacinelearning · 2026-09-08
- Yacine: GPT-6 Astra Gives Up on Unseen Tasks Out of Pure Dread — yacinelearning · 2026-09-08
- Grok Build ships triple daily updates: first-party MCP server, persistent subagents push toward full agent workspace — elonmusk · 2026-09-08
- GPT-6 Astra autonomously rebuilds 20.8km Nürburgring in Blender with 65,000 trees — reach_vb · 2026-09-08
- How GPT-6 Astra's computer use loop powers its viral Blender 3D world generation — iamrobotbear · 2026-09-08
- Fable 5.1 vs GPT-6 Astra: user calls the results very impressive — Salmaaboukarr · 2026-09-08