New paper says security agents should be judged by cost, not just success rate

Paul Kassianik · hf · 2026-07-21

This paper argues that security-agent benchmarks should not report only success rate: they should measure the full cost of reasoning, tool calls, telemetry queries, and enrichment requests.

The authors evaluate offensive and defensive security agents on Cybench and Splunk BOTS v1, comparing models at fixed cost levels rather than only best-case performance. Their main findings:

An interactive results site is available at evals.frontier.security.

Original post →

More from Research

Research channel →