Artificial Analysis Ships Intelligence Index v4.2 With Private Test Sets to Curb Gaming
burny_tech · x · 2026-09-07
Artificial Analysis announced Intelligence Index v4.2, accelerating parts of its upcoming v5 release. Changelog: new AA-Briefcase agentic knowledge-work eval with a private test set, Surge AI's GDP.pdf long-context reasoning across 4,592 PDF pages, removal of saturated GPQA Diamond, plus heavier weighting on held-out test sets to prevent gaming and upgraded grading infrastructure. Reposter Rohan joked about uneven bar heights across models.
More from Models
- Astra Solves Mensa Puzzle on Second Try While Sol Fails Twice in Reddit Test — kaljakin · 2026-09-07
- Gemini sees huge adoption outside the West, but Indian developers still find it too pricey — Aizkmusic · 2026-09-07
- After 4 Months of Failed Attempts, Astra Fixes Laptop Fan Curve in 5 Minutes — anom604 · 2026-09-07
- GPT-6 Astra can annotate baseball seams — but only on cropped ROI in ideal conditions — nickbaumann_ · 2026-09-07
- Earlier vision injection wins: why all future LLMs will be VLAs — kamalgupta09 · 2026-09-07
- Claude Opus 5 tops every major benchmark but bottoms out on user experience — gerardsans · 2026-09-07