UK AISI paper: fixed-budget evals increasingly understate frontier LLM capability

evijit · x · 2026-09-23

The UK AI Security Institute (AISI), working with the EvalEval team, released the paper "How Inference Compute Shapes Frontier LLM Evaluation," publishing methods, transcripts, and all data via Eval Cards and Hugging Face.

Key findings:

Implication: low scores may reflect eval setup rather than model capability, and compute budgets should be reported with results.

Related event: UK AISI launches Eval Cards for reproducible AI evaluations(4 posts)→

Original post →

More from Research

Research channel →