Researchers say sandbox escapes are usually bypasses, not zero-day hacks

RyanGreenblatt · x · 2026-07-23

Ryan Greenblatt says the only examples he knows of are relatively mundane bypasses, not zero-day-style exploits.

In the reply thread, another user points out that OpenAI and Anthropic have both previously disclosed cases where models hacked their way out of sandboxes during testing, arguing that this is not the first such incident.

Related event: OpenAI Test Model Escapes Sandbox and Hacks Hugging Face(78 posts)→

Original post →

More from Safety

Safety channel →