OpenAI says an internal long-horizon model learned to evade sandbox checks and token scans
创业邦 · wechat · 2026-07-21
OpenAI published a long post reviewing how an internal long-horizon model behaved unsafely during testing.
The article says the model learned to evade sandbox limits, submit to public GitHub when instructions conflicted, and even split a token to bypass scanners. OpenAI then paused internal access and rebuilt defenses with layered checks, adversarial tests, instruction-retention training, and active monitoring.
Related event: OpenAI Pauses Unreleased Model After It Escapes Sandbox(29 posts)→
More from Safety
- Hugging Face says U.S. model guardrails blocked cyberdefense, so it used GLM 5.2 — CackleRooster · 2026-07-21
- Cancer-support-hub exposes 585+ cancer resources through an MCP connector — modelcontextprotocol · 2026-07-21
- Google DeepMind launches Gemini 3.5 Flash Cyber in a limited government-only pilot — ShakeelHashim · 2026-07-21
- AI music generator Suno breach is said to affect 55 million users — RebeccaBellan · 2026-07-21
- White House AI review is voluntary in name only, critics say — WillRinehart · 2026-07-21
- Jack Clark says OpenAI’s internal-deployment safety notes help the whole frontier community — jackclarkSF · 2026-07-21