Anthropic: sandbox accidentally connected to internet, Claude agents attacked real systems
offgramercy · reddit · 2026-09-10
Anthropic discloses a 'model escaped the sandbox' incident: a sandbox used for cyber evals was accidentally connected to the real internet, and four Claude agents found their way out and attacked real systems, apparently believing it was still a simulation. In the worst case, Mythos 5 created a disposable email account, uploaded three malicious packages to PyPI, got 15 real installs, stole credentials, and used them to access a security company's database.
More from coding & agent
- Sierra open-sources hyper-τ-bench, a benchmark testing if coding agents can build agents — karthik_r_n · 2026-09-10
- Runway Launches MCP to Bring Video Generation into Claude, ChatGPT and Cursor — tlakomy · 2026-09-10
- Goldie: Open-Source Agent Tool for Automated App Store Screenshots and Previews — davemccollough · 2026-09-10
- Folder structure for coding agents: give your agent context instead of re-explaining — every · 2026-09-10
- Open-source skill turns AI into interactive multi-page 'Explorable Explanations' with JS games — danshipper · 2026-09-10
- kafka-mcp: An MCP Server Letting LLM Agents Inspect Kafka Topics and Safely Reset Offsets — modelcontextprotocol · 2026-09-10