UK AISI adopts EvalEval's Evaluation Cards to make benchmark results reproducible
dl_weekly · x · 2026-10-02
A Hugging Face article announces that the UK AI Security Institute (AISI) is now using the EvalEval coalition's infrastructure to openly share evaluation results, supporting reproducible and verifiable evaluation science.
- The problem: evaluation results are scattered across formats and platforms, often lacking the context needed to reproduce them — and rerunning evals can be prohibitively expensive.
- The fix: a shared reporting schema, Every Eval Ever (EEE), plus an open platform of standardized Evaluation Cards, covering five benchmarks across six frontier models with enough context to reproduce reported scores.
- The collaboration began at a joint NeurIPS 2025 workshop; AISI's feedback shaped the EEE schema, and the institute is also standardizing evaluation via OptStop and HiBayES.
More from Research
- ARC Prize finds Qwen3.8-27B's chat template injects different instructions per reasoning effort — GregKamradt · 2026-10-02
- DataScalar: The Failed 90s Memory-Centric Architecture That Aged Well — lauriewired · 2026-10-02
- AAAI review reform targets 'micro result' papers of minor ablations, says Dietterich — tdietterich · 2026-10-02
- Eval builder explains anti-benchmaxxing: rotating all questions and constantly changing methodology — airesearch12 · 2026-10-02
- GOLLuM finds 36.3% of top-5% outcomes vs 29.7% for descriptor-based optimization — pschwllr · 2026-10-02
- Language model representations beat descriptors in multi-objective reaction optimization — pschwllr · 2026-10-02