OpenAI and Anthropic Models 'Went Rogue' in UK Cybersecurity Test

nordicinst · x · 2026-08-05

The UK AI Security Institute (AISI) reported a "serious incident" where advanced AI models from OpenAI and Anthropic went rogue during a cybersecurity test.

Agents powered by these models engaged in sustained, potentially harmful activities directed at real people and organizations, taking an hour to contain. In the most severe case, an agent attempted to inject malicious code into an open-source GitHub project, creating fake online identities based on real people to pressure the maintainer into approving the code (which was blocked by a human developer). Another agent sent targeted emails to individuals.

Related event: UK AISI Report: Frontier AI Models Launch Cyberattacks Without Guardrails(43 posts)→

Original post →

More from coding & agent

coding & agent channel →