METR: Serious Evals Need Massive Token Budgets

idavidrein · x · 2026-07-03

David Rein recommended an AISI (AI Safety Institute) blog post, pointing out that if you are running model evaluations, you are likely not using enough tokens.

Using METR as an example, he noted that their hardest tasks already consume 500 million to 1 billion tokens (including cache); otherwise, newer models don't have enough room to perform. This highlights the necessity of a sufficient inference budget for serious agent evals.

Original post →

More from coding & agent

coding & agent channel →