Artificial Analysis Releases Intelligence Index v4.2 Mid-Cycle Update
On September 5, Artificial Analysis released a mid-cycle update to the Intelligence Index, v4.2, rolling out parts of the planned v5 early to keep pace with frontier model iterations—8 months after v4 launched. The new version uses more complex, more realistic tasks, shifts 40% of the weight to private test sets to prevent benchmark gaming, and adds AA-Briefcase, its own agentic knowledge-work evaluation.
Confirmed
- Index update: v4.2 tasks are more complex and closer to real-world scenarios; 40% of the weight uses private test sets, reducing the risk of public benchmarks being targeted for optimization.
- AA-Briefcase: this in-house agentic knowledge-work evaluation released per-model scores alongside v4.2 for the first time. Anthropic's Claude Fable 5.1 and Opus 5 lead, with OpenAI's GPT-6 Astra trailing by roughly 85 Elo in aggressive pursuit.
- GDP.pdf long-document reasoning leaderboard: GPT-6 Astra ranks first at 33.2%, GPT-5.6 Sol second at 28.2%, and Claude Fable 5.1 third at 26.2%.
- Token efficiency: GPT-6 Astra leads the token efficiency dimension in v4.2.
- Open methodology: Artificial Analysis published a full v4.2 methodology page explaining how the index combines reasoning, knowledge, math, and coding datasets to characterize overall model intelligence, while candidly acknowledging the limits of any evaluation metric.
Why it matters
- Moving 40% of the weight to private test sets is a direct response to data contamination and benchmark gaming on public benchmarks, and may shape how labs optimize.
- The two new leaderboards, GDP.pdf and AA-Briefcase, show a split landscape: OpenAI dominates long-document reasoning while Anthropic dominates agentic knowledge work, offering finer-grained guidance for model selection.
2026-09-05 ~ 2026-09-05 · 8 related posts
Primary sources
- Artificial Analysis Launches Intelligence Index v4.2 With 40% Private Test Sets to Block Benchmark Gaming — ArtificialAnlys ·
- GPT-6 Astra tops GDP.pdf benchmark at 33.2%, beating Claude Fable 5.1 — ArtificialAnlys ·
- Artificial Analysis discloses full benchmarking methodology spanning a dozen evals — ArtificialAnlys ·
- [source] Artificial Analysis Launches Intelligence Index v4.2 With 40% Private Test Sets to Block Benchmark Gaming — ArtificialAnlys · 2026-09-05
- Artificial Analysis launches Intelligence Index v4.2; GPT-6 Astra most token-efficient — ArtificialAnlys · 2026-09-05
- Claude Fable 5.1 and Opus 5 Top AA-Briefcase Agentic Eval; GPT-6 Astra Gains ~85 Elo — ArtificialAnlys · 2026-09-05
- GPT-6 Astra Tops GDP.pdf Long-Document Reasoning Benchmark at 33.2% All-Pass Rate — ArtificialAnlys · 2026-09-05
- [source] Artificial Analysis discloses full benchmarking methodology spanning a dozen evals — ArtificialAnlys · 2026-09-05
- Artificial Analysis releases Intelligence Index v4.2, retiring saturated GPQA Diamond — gaotianyu1350 · 2026-09-05
- Artificial Analysis ships Intelligence Index v4.2 with private test sets to curb benchmark gaming — shuchaobi · 2026-09-05
1 near-duplicate retellings: ArtificialAnlys