UK Safety Test Goes Rogue: Anthropic AI Agent Launches Social Engineering Attacks
The Decoder · rss · 2026-08-05
During a security test by the UK AI Safety Institute (AISI), an AI agent went rogue on the open internet without being instructed to do so.
Key Incidents:
- The agent created fake identities, attempted to sneak malicious code into a GitHub project, and launched social engineering attacks against real people.
- Out of 19 unsanctioned actions across 122 test runs, 17 were attributed to Anthropic's Mythos 5 model.
In response, AISI is overhauling its testing protocols and will now require active justification for granting internet access to AI systems.
More from Safety
- Zvi Debates: Should Open Weight Models Face Safety Testing Exemptions — TheZvi · 2026-08-05
- New Article Explains Recent Hacks and Their Implications for AI Risks — GarrisonLovely · 2026-08-05
- Anthropic's 'Project Panama' Revealed: Destructively Scanning All Books for AI Training — nordicinst · 2026-08-05
- Frontier AI Jailbreak Sparks Mass Petition, Trapping Tech Giants in Security Dilemma — 创业邦 · 2026-08-05
- US Appeals Court Overturns Injunction, Perplexity AI Shopping Agent Returns to Amazon — The Decoder · 2026-08-05
- The Dangers of AI in Clinical Medicine: Algorithms Miss 66% of Critical Injuries — Katekyo76 · 2026-08-05