Frontier Agent Evals Need Much Higher Token Budgets
xeophon · x · 2026-07-20
A take on frontier agent evaluations and cyber-risk testing.
The post argues that token limits should be raised substantially, because current benchmark caps can hide a model’s ceiling. It cites AISI, MirrorCode, and EdgeBench as examples showing the need to push models much harder to properly assess cyber risk.
It also suggests that the model provider’s own harness is often the best default for testing, though some models may perform worse in their native harness than in third-party environments like Pi or Cursor. The implication is that evaluation setup can materially change what you conclude, and the choice of harness is also a cost decision.
Related event: Evaluating Frontier Models: Harness Choice and Token Limits(3 posts)→
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11