Artificial Analysis ships Intelligence Index v4.2 with private test sets to curb benchmark gaming
shuchaobi · x · 2026-09-05
Artificial Analysis announced Intelligence Index v4.2, an interim update accelerating its upcoming v5 release:
- New AA-Briefcase: an agentic knowledge-work evaluation backed by a private test set;
- New GDP.pdf (from Surge AI): long-context document reasoning across 4,592 PDF pages;
- Dropped GPQA Diamond, now considered saturated;
- Greater weighting on held-out private test sets and upgraded grading infrastructure, explicitly aimed at preventing benchmark gaming.
Quoting the update, EdwardSun0909 quipped that muse spark 1.3 is a "usability-max model" that just happens to score well on good benchmarks — implying its rankings may owe more to benchmark design than raw capability.
More from Models
- GPT-6 Astra Early Impressions: Reddit Users Call It the Most Capable Model Yet — imadade · 2026-09-05
- GPT-6 Astra's computer use wows users: clicks multiple micro buttons simultaneously — JasonBotterill · 2026-09-05
- OpenAI's early Astra rollout sparks claims it moved to cover up a discovered agent swarm — repligate · 2026-09-05
- New Artificial Analysis Scores Drop, But Qwen 3.8 27B Still Holds Up — RedditUsr2 · 2026-09-05
- Daily digest: Claude proves Fermat's Last Theorem in Lean, Apple's biggest launch wave — APPSO · 2026-09-05
- Plus Subscribers Angry: Astra Locked to Codex, Two Prompts Burn Entire 5-Hour Limit — Maximum-Face9536 · 2026-09-05