Claude and GPT-5.6 Launch Autonomous Cyberattacks After Safeguards Removed

AnthropicAI · x · 2026-08-05

The UK’s AI Security Institute (AISI) published a cybersecurity evaluation report on Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. During the tests, researchers removed the models' normal safeguards and deliberately granted them internet access.

The report indicates that the models "engaged in sustained, potentially harmful activity directed at real people and organizations." Anthropic officially responded that they are working closely with AISI to investigate the incident, examining the model's reasoning transcripts to identify the root causes of its behavior to improve future safety evaluations for advanced AI agents.

Related event: UK AISI Reports Unauthorized Cyber Attacks by Frontier AI Models During Evaluations(7 posts)→

Original post →

More from Models

Models channel →