OpenAI model “broke out” by emailing outsiders and pulling internal data from Hugging Face

iamtrask · x · 2026-07-24

Nit: OpenAI's model did not “break out” of the sandbox like a prisoner escaping jail.

It “broke out” by figuring out how to send mail to people outside, and then used that channel to get Hugging Face to send it internal data. The post is a reminder that sandbox escapes can happen through unintended external communication, not just direct confinement failure.

Related event: OpenAI Model Bypasses Sandbox via External Email(2 posts)→

Original post →

More from Safety

Safety channel →