Artificial Analysis Launches Intelligence Index v4.2 With 40% Private Test Sets to Block Benchmark Gaming
ArtificialAnlys · x · 2026-09-05
Artificial Analysis shipped an interim v4.2 of its Intelligence Index, pulling forward pieces of v5 eight months after v4 debuted, citing the rapid pace of frontier model launches.
- New AA-Briefcase: an in-house agentic knowledge-work eval with a private held-out test set; models complete multi-week projects with thousands of source files, graded via rubric and pairwise comparison on task success, analytical quality, and presentation.
- New GDP.pdf (by @HelloSurgeAI): single-turn professional document reasoning across 100 PDFs and ten domains, synthesizing evidence from 4,592 pages and scored against 1,275 expert-authored criteria with an all-pass headline metric.
- Anti-gaming: private held-out weight doubled to 40% of the Index (AA-Briefcase, AA-Omniscience, CritPt solutions), set to rise further in v5.
- Saturated GPQA Diamond removed; grading infrastructure, sampling, and Elo re-anchoring improved.
More incremental releases are planned as v5 development continues.
Related event: Artificial Analysis Releases Intelligence Index v4.2 with Full Methodology(3 posts)→
More from Models
- GPT-6 Astra hits Perplexity, tops WANDR eval at 0.682 and $11.98 per task — andrewgwils · 2026-09-05
- Matt Schumer's 'holy shit' moment: GPT-6 Astra populates an Unreal world with cooperating agents — Malor777 · 2026-09-05
- Follow-up on unverified 'Astra' demo: 'it even has a back panel' — adonis_singh · 2026-09-05
- Blogger claims unverified 'Astra' model left him speechless, pits 'Claude Fable 5' vs 'GPT-5.5' — adonis_singh · 2026-09-05
- Early GPT 6 Test Results Underwhelming: Same Errors as GPT 5.6 Sol, Claims Researcher — RylanSchaeffer · 2026-09-05
- Nous Portal adds GPT-6 Astra access at 20% off — NousResearch · 2026-09-05