Artificial Analysis ships Capability Indices v1.1; Claude Fable 5.1 tops all six
ArtificialAnlys · x · 2026-09-15
Artificial Analysis releases Capability Indices v1.1, mapping ONET occupational tasks to weighted benchmarks across six indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare, Engineering, and Economics. Key updates: Agentic Tool Use added from AutomationBench-AA, Agentic Customer Interaction removed, GDP.pdf added to Long-Context Reasoning, AA-Briefcase added to Agentic Knowledge Work, Terminal-Bench updated to v4.0 for Engineering. Results: Claude Fable 5.1 (max) leads all six indices; GPT-6 Astra (max) second in four. Open weights are competitive: Kimi K3 leads in Finance (#8), Legal (#9), Economics (#7); DeepSeek V4.1 Flash in Strategy & Ops (#7); GLM-5.3 in Healthcare (#6) and Engineering (#7).
Related event: Artificial Analysis Capability Indices v1.1: Claude Leads All Six Domains(3 posts)→
More from Models
- Agent Arena: DeepSeek V4.1 Flash hits Pareto frontier at $0.06/task with +4.87% net improvement — arena · 2026-09-15
- 23 Days Without Claude Code: Dev Says Codex Works Better With OSS, Kimi K3 Unbeaten at Coding — Yuchenj_UW · 2026-09-15
- Marigold-V2 depth estimation demo trends on Hugging Face Spaces — toshas · 2026-09-15
- OpenAI has hundreds of contractors reading and rating your ChatGPT chats — The Decoder · 2026-09-15
- Hugging Face Rounds Up Which Open LLMs Are Best for On-Device Inference — NielsRogge · 2026-09-15
- User Claims Inference Provider nahcrof Serves Mismatched Models Under Kimi K3's Name — xeophon · 2026-09-15