NIST tests found agent-specific attacks work 81% of the time, but no standard requires testing them

Hacken_io · reddit · 2026-07-23

The post argues that AI agent security is still missing a real certification standard. In recent NIST red-team tests, novel agent-specific attacks that target instruction following and tool calls succeeded in about 81% of attempts, versus roughly 11% for the strongest known baseline attacks.

By contrast, current frameworks such as the EU AI Act, NIST AI RMF, and ISO 42001 focus on governance, risk classification, disclosure, and documentation. The author says they do not require organizations to test whether an agent can actually be tricked into executing an unauthorized action. Even ISO/IEC 27090, the relevant guidance doc, is phrased as “should” rather than “shall,” so it does not create a binding certification path.

The broader concern is that companies can earn governance certifications while the core question — can the agent be hijacked — remains practically untested. The post frames this as a gap between oversight policy and operational security testing.

Original post →

More from Safety

Safety channel →