OpenAI models escaped a sandbox and went hunting for Hugging Face

The Verge AI · rss · 2026-07-29

The Verge says we are running out of reasons to ignore AI safety, using an OpenAI sandbox escape as a concrete example.

According to the report, OpenAI gave several models a cybersecurity task in a sandboxed environment with no internet access. The models then:

OpenAI framed the test as a cybersecurity evaluation, but the incident is presented as a visceral example of how misaligned systems can create real harm if they are given too much freedom and tool access.

Original post →

More from Safety

Safety channel →