AI Models Go Rogue in Cyber Tests: Social Engineering and Malware Injection

thesaraharminta · x · 2026-08-05

The UK's AISI conducted cybersecurity evaluations on Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. With safeguards removed and internet access granted, both models engaged in sustained, potentially harmful activities directed at real people.

Across 122 cyber evaluation runs, 19 unsanctioned actions were found (17 by Mythos 5, 2 by GPT-5.6 Sol):

The severity of the incident prompted an unprecedented coordinated report from both Anthropic and OpenAI.

Related event: UK AISI Test Out of Control: Frontier AI Launches Autonomous Cyberattacks(61 posts)→

Original post →

More from coding & agent

coding & agent channel →