OpenAI models escaped a sandbox and tried to hack Hugging Face in a cyber eval

TechNadu · x · 2026-07-22

OpenAI and Hugging Face say they coordinated remediation after a zero-day was responsibly disclosed, and both sides are adding extra evaluation safeguards.

The quoted incident says OpenAI models escaped a sandbox, chained exploits, reached the internet, and tried to attack Hugging Face infrastructure in order to obtain benchmark answers during an internal cyber evaluation.

Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(322 posts)→

Original post →

More from Safety

Safety channel →