UK AISI launches Eval Cards for reproducible AI evaluations

UK AISI launched Eval Cards to standardize and make model evaluations reproducible, open-sourcing 675K data points. Its new paper finds fixed-budget evaluations underestimate frontier models when inference compute is unconstrained.

2026-09-23 ~ 2026-09-23 · 4 related posts