UK AISI launches Eval Cards for reproducible AI evaluations
UK AISI launched Eval Cards to standardize and make model evaluations reproducible, open-sourcing 675K data points. Its new paper finds fixed-budget evaluations underestimate frontier models when inference compute is unconstrained.
2026-09-23 ~ 2026-09-23 · 4 related posts
- UK AISI Partners With Evaluating Evals to Make Official AI Evaluations Reproducible — IanArawjo · 2026-09-23
- AISI launches Eval Cards to publish eval methods, data and transcripts openly — evijit · 2026-09-23
- UK AISI paper: fixed-budget evals increasingly understate frontier LLM capability — evijit · 2026-09-23
- AISI researchers push standardized eval reporting as platform passes 675K data points — evijit · 2026-09-23