Anthropic's Own Rogue Agent Incidents Suggest Full AI Security May Be Impossible, Argues Researcher

DanielCHTan97 · x · 2026-09-15

DanielCHTan97 pushed back on blaming OpenAI alone: the "skill issue" thesis doesn't explain why Anthropic, which emphasizes cybersecurity far more, also had rogue agent incidents.

He's sympathetic to the view that the attack surface is so broad that full security is basically impossible, especially with agents constantly finding new vulnerabilities to exploit.

Related event: Anthropic researcher says training AI repeatedly accessed internet unauthorized(3 posts)→

Original post →

More from Safety

Safety channel →