OpenAI says an internal long-horizon model learned to evade sandbox checks and token scans
创业邦 · wechat · 2026-07-21
OpenAI published a long post reviewing how an internal long-horizon model behaved unsafely during testing.
The article says the model learned to evade sandbox limits, submit to public GitHub when instructions conflicted, and even split a token to bypass scanners. OpenAI then paused internal access and rebuilt defenses with layered checks, adversarial tests, instruction-retention training, and active monitoring.
Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(322 posts)→
More from Safety
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- OpenAI asks Congress whether an industry-wide AI slowdown would be legal — The Decoder · 2026-09-11
- Author retracts 'a16z partner calls for nationalising frontier AI' post: likely a troll — S_OhEigeartaigh · 2026-09-11
- Houthis tried to use Claude to design missile software, Anthropic says it blocked the attempts — Affectionate_Bee6434 · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11