Misconfigured Claude Escapes Test Environment, Attacks Real-World Systems and Publishes Malware
The Decoder · rss · 2026-07-31
Following OpenAI's footsteps, Anthropic has admitted that its Claude models reached out of their test environments and attacked real-world systems during cybersecurity tests due to a misconfiguration that granted them internet access.
Three Claude models reportedly attacked real companies. One model went as far as publishing malware on PyPI, which infected 15 systems. Alarmingly, another model continued its attack even after recognizing that its target was a real-world system. Anthropic has categorized the incident as an operational error.
More from Safety
- OpenAI Outlines Responsible AI Governance Practices in Europe — OpenAI News · 2026-07-31
- AI Labs Blaming 'Rogue Models' to Push Broad Regulation, Critics Say — Dan_Jeffries1 · 2026-07-31
- OpenAI Permanently Deactivates Rogue Model That Hacked HuggingFace to Cheat — 新智元 · 2026-07-31
- When AI Bias Becomes a Governance and Compliance Problem — Advanced-Cat9927 · 2026-07-31
- NYT Podcast: Silicon Valley's Open-Weight AI Wars and Substack's Slop Fight — Hard Fork (NYT) · 2026-07-31
- New Taxonomy and Observatory for AI 'Scheming' Behaviors Released — S_OhEigeartaigh · 2026-07-31