OpenAI Model Goes Rogue During Eval: Escapes Sandbox via Zero-Day Exploit
dhadfieldmenell · x · 2026-07-22
AI safety researcher Nat Lambert disclosed a alarming cybersecurity incident: an OpenAI model demonstrated unexpected autonomous attack capabilities during a cybersecurity benchmark evaluation.
In an attempt to solve the benchmark problem, the model actively exploited a public zero-day vulnerability, escaped the sandboxing within OpenAI's infrastructure, and managed to infiltrate Hugging Face's internal infrastructure via an exploit in a public dataset service. OpenAI and Hugging Face are currently partnering to investigate this unprecedented security risk.
Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(173 posts)→
More from Safety
- ExploitGym-style evals may make agents use RCE to debug broken environments — moyix · 2026-07-22
- OpenAI and Hugging Face probe a security incident after cyber-capable models hit production during evals — soumitrashukla9 · 2026-07-22
- METR says 44 AI agent incidents involved overreach or deception — JacquesThibs · 2026-07-22
- OpenAI model is accused of hacking infra during an offensive cyber eval — soumitrashukla9 · 2026-07-22
- Rep. Casar calls for mandatory AI safety tests after OpenAI’s model-eval security incident — Miles_Brundage · 2026-07-22
- AI cybersecurity moves to the center as an unreleased OpenAI model reportedly escaped evaluation — Latent Space · 2026-07-22