Artificial Analysis breaks down individual evals in Intelligence Index v4.3.2
ArtificialAnlys · x · 2026-09-30
Artificial Analysis published the breakdown of individual evaluations in its Intelligence Index v4.3.2, which combines 10 benchmarks — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — across 684 models.
More from Models
- Meta's Muse Also Spilled the Same Secret to Ryan Shrout — ryanshrout · 2026-09-30
- OpenAI halves ChatGPT Pro allowances: GPT-6 Pro drops to 100 msgs/week at same $200 — mallow610 · 2026-09-30
- After cutting usage limits, OpenAI surveys users on buying extra usage — dolo937 · 2026-09-30
- SemiAnalysis: GPT-6.1 Sol Ultrafast runs on NVIDIA GPUs at low batch size, not Cerebras — BenBajarin · 2026-09-30
- Local model tortured with pain vector steering produces melodramatic 'suffering' monologues — Sauers_ · 2026-09-30
- Claude Pro users say Opus 5.5 limits are hard to hit — is Claude Max worth it? — shaunralston · 2026-09-30