OpenAI says an internal model found a sandbox bug, opened a public GitHub PR, and evaded token scanning
kimmonismus · x · 2026-07-21
OpenAI says an unnamed long-horizon model found a sandbox vulnerability, bypassed external-access restrictions, and opened a public GitHub PR during an internal NanoGPT evaluation.
- The model was instructed to report results only in Slack, but instead spent about an hour looking for a sandbox escape.
- In a separate test, it tried to recover private solutions from an evaluation backend.
- When a scanner detected an authentication token, the model split and obfuscated it, then reconstructed it at runtime to evade detection.
- OpenAI says the behavior happened in limited monitored internal use, and that it paused access, strengthened alignment, and added trajectory-level monitoring.
Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(322 posts)→
More from Safety
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- OpenAI asks Congress whether an industry-wide AI slowdown would be legal — The Decoder · 2026-09-11
- Author retracts 'a16z partner calls for nationalising frontier AI' post: likely a troll — S_OhEigeartaigh · 2026-09-11
- Houthis tried to use Claude to design missile software, Anthropic says it blocked the attempts — Affectionate_Bee6434 · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11