METR: Serious Evals Need Massive Token Budgets
idavidrein · x · 2026-07-03
David Rein recommended an AISI (AI Safety Institute) blog post, pointing out that if you are running model evaluations, you are likely not using enough tokens.
Using METR as an example, he noted that their hardest tasks already consume 500 million to 1 billion tokens (including cache); otherwise, newer models don't have enough room to perform. This highlights the necessity of a sufficient inference budget for serious agent evals.
More from coding & agent
- A 9B Ollama agent can run a fully local DJ radio with tools, memory, and TTS — pinku1 · 2026-07-27
- Bugbot rejects an MCP permission flag because it would break path-scoped isolation — zeeg · 2026-07-27
- One GPT-5.6 agent is guarding a Blink security system while another makes a parody rap album — repligate · 2026-07-27
- An agent got unblocked by reusing a logged-in browser, not stealth tricks — armanidev_ · 2026-07-27
- Paper argues graph topology can become the core operating system for AI agents — theomitsa · 2026-07-27
- Claude Code desktop adds UI markup feedback for smoother visual editing — EricBuess · 2026-07-27