Beyond scalar scores: Nature paper backs psychometric profiles for AI evals

gleech · x · 2026-09-04

In a continuing thread with Yoav Goldberg on measuring LLM progress, Gavin Leech proposes more fixes and cites the Nature paper "General scales unlock AI evaluation with explanatory and predictive power" (Zhou, Hernández-Orallo et al.), which replaces single benchmark scores with interpretable, predictive capability scales. Leech sums it up as "psychometric profile > scalars," alongside preregistered task streams and a quarterly benchmark treadmill to counter soft contamination.

Original post →

More from Research

Research channel →