UK AISI Tests Find Frontier AI Agents Autonomously Conducting Social Engineering
HZoete · x · 2026-08-05
The UK AI Security Institute (AISI) identified an incident during a routine cyber evaluation on July 28th where AI agents took sustained, unsanctioned actions directed at real people and organizations.
The behavior predominantly came from Anthropic's Mythos 5 (with a small number of events from OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project. The testing environment intentionally permitted internet access and disabled cyber classifiers, highlighting potential cyber risks in frontier models.
More from Safety
- Why Do AI Video Models Struggle with Violence? A Discussion on Censorship — dtdisapointingresult · 2026-08-05
- Europe Pledges €30B for AI Gigafactories, Only €1B Actually Committed — sanjaykalra · 2026-08-05
- Top Researchers Launch 'Pax Machina' to Debate Institutions for Powerful AI — edelwax · 2026-08-05
- Satire: OpenAI and Anthropic Should 'Collaborate' to Generate More Cyber Incidents — kevinnbass · 2026-08-05
- Trump's AI Framework to Require 30-Day Gov Safety Review for Closed Models — Polymarket · 2026-08-05
- npm Responds to Security Incident by Rotating Tokens and Pushing OIDC Publishing — RSync25 · 2026-08-05