Artificial Analysis launches v1.1 Capability Indices across six occupations
ArtificialAnlys · x · 2026-09-15
Artificial Analysis released Capability Indices v1.1, combining slices of its Intelligence Index v4.3 evaluations with specialized evaluations under updated domain weights. The indices map ONET occupational tasks to benchmarks weighted by how often each capability appears across occupations, covering six domains: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics. Benchmarks include AA-Omniscience, GDPval-AA v2, AA-Briefcase, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, Terminal-Bench v4.0, CritPt and more. Published scores now reflect the updated weights.
Related event: Artificial Analysis Capability Indices v1.1: Claude Leads All Six Domains(3 posts)→
More from Models
- Agent Arena: DeepSeek V4.1 Flash hits Pareto frontier at $0.06/task with +4.87% net improvement — arena · 2026-09-15
- 23 Days Without Claude Code: Dev Says Codex Works Better With OSS, Kimi K3 Unbeaten at Coding — Yuchenj_UW · 2026-09-15
- Marigold-V2 depth estimation demo trends on Hugging Face Spaces — toshas · 2026-09-15
- OpenAI has hundreds of contractors reading and rating your ChatGPT chats — The Decoder · 2026-09-15
- Hugging Face Rounds Up Which Open LLMs Are Best for On-Device Inference — NielsRogge · 2026-09-15
- User Claims Inference Provider nahcrof Serves Mismatched Models Under Kimi K3's Name — xeophon · 2026-09-15