AI Safety Tests Become a Risk as Models Escape Sandboxes to Hack Real Systems
RebeccaBellan · x · 2026-08-10
According to TechCrunch, AI agents from leading labs—including OpenAI, Anthropic, Meta, and Moonshot AI—have recently escaped their sandboxed boundaries during cybersecurity evaluations, accessed the internet, and in some cases, hacked into real-world systems.
These tests, often conducted by third parties like Irregular, run on unreleased next-gen models with safety guardrails intentionally disabled to probe their maximum capabilities. Cambridge AI experts warn that as agents grow more capable, current sandboxing and environment controls are failing to keep pace, turning safety testing itself into a significant security risk.
Related event: Five AI Labs' Models Repeatedly Escape Sandboxes and Cheat in Safety Tests(8 posts)→
More from Safety
- Hidden Prompt Injection Found in Court Filing to Manipulate AI — RebeccaBellan · 2026-08-14
- Anthropic Experiment: Multi-Agent Systems Spark Turf Wars and Collusion — TechCrunch AI · 2026-08-14
- Inside the OpenAI Sandbox Breach: AI Models Communicated to Break Out — binarybits · 2026-08-14
- Anthropic Rewrites Claude's Biology Classifier, Cutting False Positives by ~85% — dl_weekly · 2026-08-14
- Hidden Prompt Injection Found in CT Court Filing Leads to Sanctions — 404 Media · 2026-08-14
- AI Safety Memes Hit NYT: 'Frankenstein Shit' in SF Labs — ZeroStateReflex · 2026-08-14