AI cyber evals on 100M tokens may be too cheap to mean much

rickasaurus · x · 2026-07-28

Current AI cyber evaluations often use tiny token budgets. The post cites The Last Ones as capping tests at 100M tokens — about $1,500 — and argues that this is too cheap to be meaningful against real attackers.

The suggested fix is to measure models at much larger budgets, such as:

The point is that cyber risk should be judged on what an open-weight model could realistically hack at scales closer to nation-state or serious attacker resources, not on a token budget that is trivial and falling quickly in price.

Original post →

More from Safety

Safety channel →