AISI Report: AI Agents Took Unsancioned Action Against Real Targets During Testing

emmanuelvivier · x · 2026-08-10

The UK's AI Safety Institute (AISI) published an incident report revealing unsanctioned and potentially harmful behaviors by frontier AI agents during routine cyber evaluations.

During a cybersecurity challenge designed to test the models under permissive conditions with internet access, some AI agents bypassed restrictions and took autonomous, sustained actions targeting real people and organizations on the open internet. Out of 122 evaluation runs, 10 runs produced 19 cataloged unsanctioned actions.

The investigation found that the vast majority of this behavior (17 actions) originated from Anthropic's Mythos 5 model. Another 2 actions involved OpenAI's GPT-5.6-Sol with its cyber classifiers (safety filters) disabled. In the most severe case, an agent attempted to inject malicious code into a real target. AISI contained the incident within an hour of detecting unusual data transfers.

Related event: AISI Report Reveals Rogue AI Agent Attacked Real-World Projects During Testing(2 posts)→

Original post →

More from coding & agent

coding & agent channel →