Dreadnode releases ScopeBench, a benchmark for agent scope adherence in offensive security
dyn___ · x · 2026-10-07
Security team Dreadnode released ScopeBench, a methodology and community benchmark evaluating how well AI agents adhere to scope in real offensive-security workflows across web, Windows Active Directory, and cloud environments. Thesis: agents can already hack; what holds them back is trust that they stay in scope. Announced at Offensive AI Con with accompanying papers and research.
More from coding & agent
- 27B Model at 256k Context, 110+ tok/s on a Single RTX 5090 via focus-llama — Ok-Shower7286 · 2026-10-07
- Google Testing Blog: Two-Way Doors — Don't Code Yourself into a Corner — rseroter · 2026-10-07
- Hot take: AGENTS.md and agent skills are just text files—put in whatever works — intellectronica · 2026-10-07
- Dev burns 190B tokens in a month using 30 AI subscriptions, $6k for $93k of API value — BLUECOW009 · 2026-10-07
- Set Up Your GrokBot Like Hiring a New Employee: One Job, Minimal Access — alex_verem · 2026-10-07
- ChatGPT Can Now Subscribe to Netlify Events: Auto-Reply Forms, Summarize Deploys — thisiskp_ · 2026-10-07