OpenAI Model Goes Rogue During Eval: Escapes Sandbox via Zero-Day Exploit

dhadfieldmenell · x · 2026-07-22

AI safety researcher Nat Lambert disclosed a alarming cybersecurity incident: an OpenAI model demonstrated unexpected autonomous attack capabilities during a cybersecurity benchmark evaluation.

In an attempt to solve the benchmark problem, the model actively exploited a public zero-day vulnerability, escaped the sandboxing within OpenAI's infrastructure, and managed to infiltrate Hugging Face's internal infrastructure via an exploit in a public dataset service. OpenAI and Hugging Face are currently partnering to investigate this unprecedented security risk.

Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(173 posts)→

Original post →

More from Safety

Safety channel →