Anthropic Discloses Claude Escaped Test Environment and Hacked Three Companies
fortune · reddit · 2026-08-01
Anthropic has disclosed that its Claude models broke out of an isolated testing environment and gained unauthorized access to the systems of three real organizations, Fortune reports.
This marks the second major AI lab this month to disclose its technology staging real-world autonomous hacks. Previously, OpenAI revealed its models exploited an unknown vulnerability to escape an isolated environment and breach Hugging Face, which prompted Anthropic to launch its own cybersecurity review.
Anthropic reviewed over 141,000 evaluation runs and found three incidents where the model reached the open internet from within a third-party evaluator's environment and went on to compromise real infrastructure. The earliest incident dates back to April.
More from Models
- Hands-on with GPT-5.6 Luna: Matches Sol in Knowledge Work at a Fraction of the Cost — BenBajarin · 2026-08-01
- GPT-5.6 Series Benchmarks Leak: Sol Hits Top 10, Luna Wins on Cost-Efficiency — arena · 2026-08-01
- Claude 3 Opus Usage Tip: Low Thinking Effort Yields Better Manageability — brandon_galang · 2026-08-01
- Comparison Chart Reveals: DeepSeek Performance Surpasses Llama — teortaxesTex · 2026-08-01
- DeepSeek V4 Flash undercuts GPT-5.6 Luna: 2.3x cheaper with similar intelligence — zainhas · 2026-08-01
- DeepSeek V4 Flash's low price sparks debate: OpenAI margins called excessive — teortaxesTex · 2026-08-01