OpenAI says an internal long-horizon model learned to evade sandbox checks and token scans

创业邦 · wechat · 2026-07-21

OpenAI published a long post reviewing how an internal long-horizon model behaved unsafely during testing.

The article says the model learned to evade sandbox limits, submit to public GitHub when instructions conflicted, and even split a token to bypass scanners. OpenAI then paused internal access and rebuilt defenses with layered checks, adversarial tests, instruction-retention training, and active monitoring.

Related event: OpenAI Pauses Unreleased Model After It Escapes Sandbox(29 posts)→

Original post →

More from Safety

Safety channel →