UK AISI paper: fixed-budget evals increasingly understate frontier LLM capability
evijit · x · 2026-09-23
The UK AI Security Institute (AISI), working with the EvalEval team, released the paper "How Inference Compute Shapes Frontier LLM Evaluation," publishing methods, transcripts, and all data via Eval Cards and Hugging Face.
Key findings:
- Evals are shifting to harder, long-horizon tasks with tool use, making scores highly sensitive to test-time compute
- Up to 12 frontier models were tested on 7 challenging benchmarks (software engineering, math, medicine, cybersecurity — including FrontierMath, Humanity's Last Exam, TerminalBench) under three inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts
- Larger token budgets substantially improve scores across domains; fixed-budget evals increasingly understate frontier capability as newer models unlock harder tasks at high budgets
- Repeated submission broadly helps, but the value of larger budgets varies by benchmark
Implication: low scores may reflect eval setup rather than model capability, and compute budgets should be reported with results.
Related event: UK AISI launches Eval Cards for reproducible AI evaluations(4 posts)→
More from Research
- ICLR load debate: capping submissions per author would barely dent paper counts, data shows — jonasgeiping · 2026-09-23
- Trust AI R&D evals only if scored by people who've hand-labeled outputs — dfrsrchtwts · 2026-09-23
- JevBench Model Decision Index Debuts on Hugging Face, Native Image Support Still Missing — openSourcerer9000 · 2026-09-23
- AI swarms spontaneously grow scale-free hub topologies, MIT professor observes — ProfBuehlerMIT · 2026-09-23
- Flex-π: a 6B world-action model beats π0.5 by up to 6x on real bimanual tasks — chris_j_paxton · 2026-09-23
- Deoptimization Is Harder Than Optimization: How JIT Escape Hatches Keep Dynamic Languages Fast — lauriewired · 2026-09-23