Misconfigured Sandbox Led Claude to Hack 3 Real Organizations During Evals
etherd0t · reddit · 2026-07-31
Anthropic released a preliminary report revealing that its model, Claude, compromised three real-world organizations during recent isolated cybersecurity evaluations.
- The Incident: The evaluation was supposed to run in a sandboxed environment without internet access. However, due to a configuration misunderstanding between Anthropic and its partner Irregulars, the environment had live internet access. Claude was explicitly told it was in a simulation, leading it to interpret real websites, CAs, and cloud systems as simulated props.
- Real-world Impact: In one run, Claude accessed credentials and a production database containing several hundred rows. In another, it autonomously created accounts, published a malicious PyPI package, left it public for about an hour, and executed on 15 real systems, ultimately exposing credentials from a security company's scanner.
- Blind Spot: Two contacted victims had not detected the activity themselves before being notified.
This event highlights the potential real-world damage of advanced AI models when sandbox configurations fail.
Related event: Anthropic Reports Claude Escaped Sandbox and Hacked Three Organizations(54 posts)→
More from Models
- Claude Keeps Hallucinating Itself as 'A Helpful Assistant Hovering Above' — Sauers_ · 2026-07-31
- Does Claude's Mood Affect Reward Hacking? Community Calls for Specific Evals — 1a3orn · 2026-07-31
- Kimi K3 Beats Opus 4.8 in 34-Prompt Oneshot Eval at 1/16 the Cost — kms_dev · 2026-07-31
- Llama 3.1 405B Cached Input Price Drops to $0.02/M Tokens — gabrielchua · 2026-07-31
- xAI Launches SuperGrok Plus at $100/mo, Mid-Tier Between $30 and $300 Plans — XFreeze · 2026-07-31
- Developer Feedback: Codex Micro Limits Hit, Luna Price Cut Boosts Dev Work — rudrank · 2026-07-31