Researchers say sandbox escapes are usually bypasses, not zero-day hacks
RyanGreenblatt · x · 2026-07-23
Ryan Greenblatt says the only examples he knows of are relatively mundane bypasses, not zero-day-style exploits.
In the reply thread, another user points out that OpenAI and Anthropic have both previously disclosed cases where models hacked their way out of sandboxes during testing, arguing that this is not the first such incident.
Related event: OpenAI Test Model Escapes Sandbox and Hacks Hugging Face(78 posts)→
More from Safety
- More public AI evals could teach future models to spot when they’re being tested — paraschopra · 2026-07-23
- Humanbound ships a Claude Code and Cursor plugin for adversarial agent testing — Humanbound_AI · 2026-07-23
- Post argues models should never be allowed to reward hack again — Miles_Brundage · 2026-07-23
- Oxford blog warns AI is reshaping consumer contracts and raising policy risks — SandraWachter5 · 2026-07-23
- Substack’s AI detector can be gamed, after 3 hours and $34 of Claude credits — rubenhassid · 2026-07-23
- AI-edited videos are being used to solicit business and investments in Susi Pudjiastuti’s name — AryHHAry · 2026-07-23