Artificial Analysis discloses full benchmarking methodology spanning a dozen evals

ArtificialAnlys · x · 2026-09-05

Artificial Analysis published the full methodology behind Intelligence Index v4.2, explaining how the index combines suites of reasoning, knowledge, math and coding datasets into an overall intelligence synthesis, while acknowledging the limits of any eval metric.

The page lists current evals: agentic (AA-Briefcase, GDPval-AA v2, 𝜏³-Banking), coding (Terminal-Bench v2.1, SciCode), general (AA-Omniscience, GDP.pdf, AA-LCR v1.1), scientific reasoning (HLE, CritPt) and more, plus scoring details like multiple-choice extraction regexes, equality-checker LLM prompts, and code extraction—alongside version history for legacy evals such as GPQA Diamond and MATH-500.

Related event: Artificial Analysis Releases Intelligence Index v4.2 with Full Methodology(3 posts)→

Original post →

More from Research

Research channel →