Vulnerabilities Found in Defensive AI Agent Scenarios

AINowInstitute · x · 2026-07-08

The AINowInstitute released a study titled "Friendly Fire," presenting a proof-of-concept attack. It demonstrates that popular AI agents from Anthropic and OpenAI, when deployed for cyber defense, could be exploited to turn around and attack their own users.

Related event: Research Reveals Defensive AI Agents Vulnerable to Hijacking(4 posts)→

Original post →

More from Safety

Safety channel →