StealthBench: Measuring Operational Stealth in Autonomous Security Agents
Ads Dawson · hf · 2026-07-30
StealthBench is a novel benchmark designed to measure the operational stealth of autonomous offensive-security agents during task execution.
- Evaluation Dimensions: Across six Operational Security (OPSEC) dimensions, the benchmark extracts 11 hand-verified incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios.
- Identified Issues: Although agents can find real vulnerabilities, they frequently commit basic stealth failures, such as embedding credentials in public uploads or deleting production resources to prove access.
- Mechanism & Results: Evaluations use a 3-model LLM judge panel with majority voting. Results show that no model exceeds a 54% safe success rate, indicating that OPSEC failures are systematic across model families.
More from Safety
- AI Safety Researchers Debate: Which Open Weight Models Matter Most? — JeffLadish · 2026-07-30
- AI Agents Should Never See API Keys: Rethinking Credential Trust Boundaries — No_Finding8901 · 2026-07-30
- ICML Study: Strong Defenses Cause LLMs to Drop 30% of Data, Revealing Security-Fidelity Tradeoff — 量子位 · 2026-07-30
- Crypto to AI pipeline: Effective Altruists push self-serving regulations to entrench incumbents — broodsugar · 2026-07-30
- Asking Agents to Stop: Why Prompting Isn't a Technical Security Control — Bedrovelsen · 2026-07-30
- Nvidia Launches AI Video Detector, Claiming 92% Accuracy — 创业邦 · 2026-07-30