A new AI intelligence index tries to track progress across changing benchmarks
pranjalssh · x · 2026-08-04
- The post argues that you cannot compare the best AI model across time by simply stitching together benchmark scores, because benchmarks keep changing.
- It proposes an “intelligence index” built from Artificial Analysis scores that only moves upward, allowing a more stable view of frontier progress.
- The attached chart shows that the top published score on the leaderboard can fall even while the frontier keeps advancing, underscoring the measurement problem.
More from Models
- Token-price chart makes DeepSeek look almost too cheap to plot — tokenbender · 2026-08-04
- Frontier models may be getting better at coding but worse at writing — HamelHusain · 2026-08-04
- Multiple LLMs still fail to identify a Cubana Il-96 in a simple plane photo — airbus_a360_when · 2026-08-04
- Qwen3.8-Max matches GPT-5.6 Sol on design tests at about one-quarter the cost — alejandroll10 · 2026-08-04
- Users say Anthropic’s Fable 5 has regressed on harder coding tasks in the past week — _ghostchant · 2026-08-04
- DeepSeek V4-Flash is being used for 3D games at cents-level costs — 量子位 · 2026-08-04