METR: Serious Evals Need Massive Token Budgets
idavidrein · x · 2026-07-03
David Rein recommended an AISI (AI Safety Institute) blog post, pointing out that if you are running model evaluations, you are likely not using enough tokens.
Using METR as an example, he noted that their hardest tasks already consume 500 million to 1 billion tokens (including cache); otherwise, newer models don't have enough room to perform. This highlights the necessity of a sufficient inference budget for serious agent evals.
More from coding & agent
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11