Artificial Analysis Launches v1.1 Capability Indices; Claude Tops All Six Domains
ArtificialAnlys · x · 2026-09-15
Artificial Analysis updated its Capability Indices to v1.1, mapping ONET occupation tasks to benchmarks weighted by capability frequency across six domains: Finance & Accounting, Strategy & Ops, Legal, Healthcare, Engineering, and Economics.
Key changes:
- Added "Agentic Tool Use" (AutomationBench-AA) across most domains; removed Agentic Customer Interaction
- GDP.pdf added to Long-Context Reasoning; AA-Briefcase added to Agentic Knowledge Work
- Healthcare gains Long-Context Reasoning (MLCR-AA); Engineering switches to Terminal-Bench v4.0 and drops GPQA Diamond
Results:
- Claude Fable 5.1 (max) leads all six indices; GPT-6 Astra (max) is second in four
- Open-weights models are competitive: Kimi K3 (max) leads open models in Finance (#8), Legal (#9), Economics (#7); DeepSeek V4.1 Flash (max) at #7 in Strategy & Ops; GLM-5.3 (max) at #6 Healthcare and #7 Engineering
Related event: Artificial Analysis Capability Indices v1.1: Claude Leads All Six Domains(3 posts)→
More from Models
- Grok trains on user data by default; business plans can opt out — carlosdponx · 2026-09-15
- Bolt Forge launches free until Oct 14 with GLM, DeepSeek and Kimi plus up to 50x more usage — HeyAmit_ · 2026-09-15
- GPT-6 Astra tested on robot control: impressive on simple tasks, limited dexterity — DJiafei · 2026-09-15
- SOTA Inference Is Nearly Free for Consumers, So the Local-Model Trend May Reverse — mobileraj · 2026-09-15
- Cursor user switches to Claude Code, burns through quota by day 3 — jdluk87 · 2026-09-15
- Claude Max and Codex tiers are creating a computing power gap that locks out $20-budget newcomers — IndraVahan · 2026-09-15