New paper says security agents should be judged by cost, not just success rate
Paul Kassianik · hf · 2026-07-21
This paper argues that security-agent benchmarks should not report only success rate: they should measure the full cost of reasoning, tool calls, telemetry queries, and enrichment requests.
The authors evaluate offensive and defensive security agents on Cybench and Splunk BOTS v1, comparing models at fixed cost levels rather than only best-case performance. Their main findings:
- Offensive CTF tasks improve with more test-time compute, and scaled open-weight models can get close to frontier proprietary systems while staying cost-competitive.
- Defensive SOC investigations do not scale the same way; success depends more on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget.
- They propose cost-aware, SOC-native evaluations as a clearer way to judge which models are practically useful in security workflows.
An interactive results site is available at evals.frontier.security.
More from Research
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- Chinese AI labs are now treating distillation obfuscation as the top research topic — pmddomingos · 2026-07-22
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22