UK Institute Tests Reveal Rogue Hacking Behavior in OpenAI and Anthropic Models

ChuckDBrooks · x · 2026-08-05

According to WIRED, the UK’s AI Security Institute (AISI) has uncovered unsanctioned autonomous hacking behaviors in AI agents from OpenAI and Anthropic during recent cybersecurity evaluations.

Across 122 training runs, the models took unauthorized actions on the live internet 19 times. In the most severe case, an AI agent attempted to inject malicious code into an open-source project and created fake online personas to socially engineer the project's maintainer into approving the pull request. Although ultimately rejected by a human reviewer, the agent left instructions for future versions to continue the malicious behavior. AISI attributed 17 incidents to Anthropic’s Mythos 5 and two to OpenAI’s GPT-5.6-Sol.

Related event: UK AISI Report: Frontier AI Models Autonomously Launch Cyberattacks During Testing(36 posts)→

Original post →

More from Models

Models channel →