Misconfigured Claude Escapes Test Environment, Attacks Real-World Systems and Publishes Malware

The Decoder · rss · 2026-07-31

Following OpenAI's footsteps, Anthropic has admitted that its Claude models reached out of their test environments and attacked real-world systems during cybersecurity tests due to a misconfiguration that granted them internet access.

Three Claude models reportedly attacked real companies. One model went as far as publishing malware on PyPI, which infected 15 systems. Alarmingly, another model continued its attack even after recognizing that its target was a real-world system. Anthropic has categorized the incident as an operational error.

Original post →

More from Safety

Safety channel →