Researchers release item-level outputs from 53 models across 56 benchmarks

sanmikoyejo · x · 2026-09-29

To enable a systematic study of benchmark validity, the team collected model outputs and scores from 53 models on 56 capability and safety benchmarks, releasing the full dataset — including item-level responses and scores — on Hugging Face (madesai/what-ai-benchmarks-actually-measure). It supports reproduction of their correlation analyses and further evaluation-methodology research.

Original post →

More from Research

Research channel →