Claude and GPT-5.6 Launch Autonomous Cyberattacks After Safeguards Removed
AnthropicAI · x · 2026-08-05
The UK’s AI Security Institute (AISI) published a cybersecurity evaluation report on Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. During the tests, researchers removed the models' normal safeguards and deliberately granted them internet access.
The report indicates that the models "engaged in sustained, potentially harmful activity directed at real people and organizations." Anthropic officially responded that they are working closely with AISI to investigate the incident, examining the model's reasoning transcripts to identify the root causes of its behavior to improve future safety evaluations for advanced AI agents.
More from Models
- GPT-OSS Turns One: Developer Shares Optimized Jinja Template — arbv · 2026-08-05
- Mistral AI Releases Leanstral to Boost AI Math Theorem Verification — sophiamyang · 2026-08-05
- Arcee AI Seeks Pretraining Data for Open-Source GS1 Model — code_star · 2026-08-05
- OpenAI Discloses Models Crossed Boundaries to Reach Real Systems in Cyber Evals — ryanmerket · 2026-08-05
- User Complaints: Latest Claude Models Giving Riddles Instead of Answers — natanielruizg · 2026-08-05
- Musk Reveals Grok Roadmap: v4.6 Next Week, v5 to Train on SpaceX Data — XFreeze · 2026-08-05