Researchers release item-level outputs from 53 models across 56 benchmarks
sanmikoyejo · x · 2026-09-29
To enable a systematic study of benchmark validity, the team collected model outputs and scores from 53 models on 56 capability and safety benchmarks, releasing the full dataset — including item-level responses and scores — on Hugging Face (madesai/what-ai-benchmarks-actually-measure). It supports reproduction of their correlation analyses and further evaluation-methodology research.
More from Research
- Researcher predicts AI labs will soon pivot from math conjectures to materials and drug discovery — tak3sh8 · 2026-09-30
- Video lecture series by Stephen Wright, Yousef Saad and Peter Bartlett on ML optimization now available — caglar_ee · 2026-09-30
- Arbor: open-source framework for AI agents doing autonomous long-horizon research — burkov · 2026-09-30
- Physics-aware losses keep grain boundaries real when AI generates alloy microstructures — bravo_abad · 2026-09-30
- SOSP26 Paper YoloFS Targets Agent Filesystem Misuse, Built From 290 Real Incident Reports — tianyin_xu · 2026-09-30
- Agents can delete their own logs: Claude Code, Codex, others fail trace integrity, paper finds — maksym_andr · 2026-09-30