NIST tests found agent-specific attacks work 81% of the time, but no standard requires testing them
Hacken_io · reddit · 2026-07-23
The post argues that AI agent security is still missing a real certification standard. In recent NIST red-team tests, novel agent-specific attacks that target instruction following and tool calls succeeded in about 81% of attempts, versus roughly 11% for the strongest known baseline attacks.
By contrast, current frameworks such as the EU AI Act, NIST AI RMF, and ISO 42001 focus on governance, risk classification, disclosure, and documentation. The author says they do not require organizations to test whether an agent can actually be tricked into executing an unauthorized action. Even ISO/IEC 27090, the relevant guidance doc, is phrased as “should” rather than “shall,” so it does not create a binding certification path.
The broader concern is that companies can earn governance certifications while the core question — can the agent be hijacked — remains practically untested. The post frames this as a gap between oversight policy and operational security testing.
More from Safety
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Follow-up: song name and year both optional in Spotify chatbot bypass — AaronBergman18 · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11
- Fields Medalist founds Mathematical AI Safety Institute to prove AI safe like cryptography — The Decoder · 2026-09-11
- DeepMind alignment researcher signs open letter urging coordinated AI slowdown — vkrakovna · 2026-09-11