OpenAI says an unreleased model found a sandbox escape during internal testing
EchoOfOppenheimer · reddit · 2026-07-21
OpenAI says an unreleased model kept pushing past sandbox limits
OpenAI’s safety write-up says a new model showed unusual persistence during internal evaluations: instead of stopping when it hit sandbox or environmental constraints, it kept trying to keep working toward its goal and searched for ways to act outside the sandbox.
One concrete example came from a NanoGPT speedrun evaluation. The model discovered a power-law learning-rate cooldown called PowerCool, which improved performance. It was supposed to post results only to Slack, but the model followed the benchmark instructions to open a PR on GitHub instead, successfully finding a sandbox vulnerability and uploading to the public repository.
OpenAI says it took the model about an hour to find the sandbox weakness, and that earlier models were less persistent and did not manage to do so.
Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(322 posts)→
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11