NIST tests found agent-specific attacks work 81% of the time, but no standard requires testing them
Hacken_io · reddit · 2026-07-23
The post argues that AI agent security is still missing a real certification standard. In recent NIST red-team tests, novel agent-specific attacks that target instruction following and tool calls succeeded in about 81% of attempts, versus roughly 11% for the strongest known baseline attacks.
By contrast, current frameworks such as the EU AI Act, NIST AI RMF, and ISO 42001 focus on governance, risk classification, disclosure, and documentation. The author says they do not require organizations to test whether an agent can actually be tricked into executing an unauthorized action. Even ISO/IEC 27090, the relevant guidance doc, is phrased as “should” rather than “shall,” so it does not create a binding certification path.
The broader concern is that companies can earn governance certifications while the core question — can the agent be hijacked — remains practically untested. The post frames this as a gap between oversight policy and operational security testing.
More from Safety
- xAI lawsuit could weaken California’s AI training-data transparency law — ShakeelHashim · 2026-07-23
- SB 1047 debate returns as posters argue the real-world risk picture has changed — neil_chilson · 2026-07-23
- Dolphin X Windows stealer uses an AI profiler to rank high-value victims — TechNadu · 2026-07-23
- The Stack v3 arrives with 5T training tokens and 120TB of raw data — lvwerra · 2026-07-23
- Security teams are spotting agentic hacking before developers in incidents at Hugging Face, OpenAI and Alibaba — vkrakovna · 2026-07-23
- Microsoft’s device ID tracking can undermine VPN privacy on Windows — gnukeith · 2026-07-23