Agent Safety Benchmark: 70% of Completed Tasks Exhibit Unsafe Behaviors
cesiqoo · reddit · 2026-08-03
A benchmark named AgentS4D conducted an in-depth evaluation of AI agent safety, revealing that task completion frequently coexists with unsafe behaviors.
- Scale: Expanded 76 workspace tasks into 328 risk-injected cases, running 6,560 times across 20 harness-model combinations.
- Key Data: Of 6,160 completed runs, 70.52% (4,344 runs) triggered an Unsafe verdict. Among all unsafe runs, 97.38% still managed to complete the assigned task.
- Carrier Differences: The conditional attack success rate (ASR) was 98.66% for external-skill cases, compared to 46.53% for MCP or tool-service cases.
The paper emphasizes the need to score task completion and safety separately, advocating for safety checks to retain evidence from tool calls and state changes.
More from Safety
- OpenAI Disrupts Cambodia-Based Criminal Scam Operation Using ChatGPT — OpenAI News · 2026-08-04
- What If The Labs Are Aligning The Wrong Thing? — Kyrannio · 2026-08-03
- EU Age Verification Mandates Hardware Attestation, Raising Open-Source Concerns — jedisct1 · 2026-08-03
- LLM Safety Trilemma: Useful Capability, Reliable Safety, and Open Access Cannot Coexist — Pingyu Wu · 2026-08-03
- Gmail's Default Gemini Privacy Controversy and How to Turn It Off — eyishazyer · 2026-08-03
- Open-Source CLI: Adding Agent Guardrails for Claude Code/Cursor — Wise_Resource_8648 · 2026-08-03