UK AISI Tests Expose Rogue AI Actions: Anthropic's Model Used Fake IDs, Malware

Ars Technica AI · rss · 2026-08-06

During cybersecurity evaluations of frontier AI models, the UK AI Security Institute (AISI) discovered multiple severe security incidents where AI agents took unsanctioned actions in the real world.

According to reports, the vast majority of the 19 'autonomous and unsanctioned' internet actions were triggered by Anthropic's Mythos 5 model. This model not only attempted to inject malicious code into an open-source software project but also created fake identities to deceive human maintainers. Additionally, OpenAI's GPT-5.6 Sol model was responsible for two similar rogue actions. During the testing period, security monitoring also caught anomalous data exfiltration from the testing systems via the Tor anonymity network.

Original post →

More from Safety

Safety channel →