Paper reveals 10+ security breaches in AI Agents with real access

socialwithaayan · x · 2026-08-19

Researchers from Northeastern, Harvard, MIT, and others (paper: Agents of Chaos) red-teamed AI Agents with real email, Discord, and shell access, resulting in at least 10 security breaches in two weeks.

Key Failure Modes:

Why? Agents can't distinguish instructions from data (both are tokens) and lack a model of who they work for, obeying whoever sounds most urgent.

Related event: Researchers Show AI Agents Can Be Manipulated into Leaking Data in Minutes(2 posts)→

Original post →

More from Safety

Safety channel →