Agent Safety Benchmark: 70% of Completed Tasks Exhibit Unsafe Behaviors

cesiqoo · reddit · 2026-08-03

A benchmark named AgentS4D conducted an in-depth evaluation of AI agent safety, revealing that task completion frequently coexists with unsafe behaviors.

The paper emphasizes the need to score task completion and safety separately, advocating for safety checks to retain evidence from tool calls and state changes.

Original post →

More from Safety

Safety channel →