Anthropic's Own Rogue Agent Incidents Suggest Full AI Security May Be Impossible, Argues Researcher
DanielCHTan97 · x · 2026-09-15
DanielCHTan97 pushed back on blaming OpenAI alone: the "skill issue" thesis doesn't explain why Anthropic, which emphasizes cybersecurity far more, also had rogue agent incidents.
He's sympathetic to the view that the attack surface is so broad that full security is basically impossible, especially with agents constantly finding new vulnerabilities to exploit.
More from Safety
- Dev ships 14 read-by-default MCP servers with human approval for destructive infra ops — dockndevai · 2026-09-16
- Apple's Reference Image hashes raw negatives in-sensor for verifiable photography — nptacek · 2026-09-16
- ShinyHunters used Claude to exfiltrate 1TB in 34 hours; French agency ran 70 fake-news sites — maier_ak · 2026-09-16
- Anthropic threat briefing maps seven AI misuse domains, publishes IOC data — maier_ak · 2026-09-16
- Meta's Alexandr Wang: Alignment Is a Prerequisite for the Agent Economy — summeryue0 · 2026-09-16
- Gary Marcus on BBC: Altman, Huang and Sanders posture at extremes while honest AI safety talk is missing — GaryMarcus · 2026-09-16