Frontier Agent Evals Need Much Higher Token Budgets

xeophon · x · 2026-07-20

A take on frontier agent evaluations and cyber-risk testing.

The post argues that token limits should be raised substantially, because current benchmark caps can hide a model’s ceiling. It cites AISI, MirrorCode, and EdgeBench as examples showing the need to push models much harder to properly assess cyber risk.

It also suggests that the model provider’s own harness is often the best default for testing, though some models may perform worse in their native harness than in third-party environments like Pi or Cursor. The implication is that evaluation setup can materially change what you conclude, and the choice of harness is also a cost decision.

Related event: Evaluating Frontier Models: Harness Choice and Token Limits(3 posts)→

Original post →

More from Safety

Safety channel →