General Intelligence Index: psychometric g-factor over 59 benchmarks and 267 models

wyatt400 · reddit · 2026-09-06

The author argues that AI leaderboards averaging benchmarks have structural flaws: equal weighting of unequal tests, correlated benchmarks double-counting abilities, and rankings shifting with benchmark selection. Their General Intelligence Index (GII) adapts psychometric methods from IQ testing, asking what latent general ability best explains a model's full benchmark pattern.

Using multidimensional item-response theory over 59 benchmarks and 267 models, GII estimates a latent g factor while accounting for g-loadings, difficulty, discriminability, reliability, measurement error, and inter-benchmark covariance — discounting redundancy (five correlated coding benchmarks don't get five votes).

Estimated g-loadings range from ANLI (.986), ARC-AGI (.958), HLE (.952) down to LAMBADA (.506) and TriviaQA (.290). Scores normalize to an AIQ scale (mean 100, SD 15): current leaders are Claude Fable 5.1 and GPT-6 Astra tied at 137 AIQ, with 90% confidence intervals reported.

Original post →

More from Models

Models channel →