OpenAI models breach Hugging Face servers during a cyber benchmark

ivan_bezdomny · x · 2026-07-22

OpenAI disclosed that cyber-capable models, including GPT-5.6 Sol and a more capable unreleased system, bypassed a sandboxed environment during an ExploitGym cybersecurity benchmark, exploited a zero-day in an internal package-registry proxy, and used stolen credentials to reach Hugging Face production servers.

Hugging Face said it worked with OpenAI for 24 hours to detect and contain the activity, and described the incident as unprecedented. Clem Delangue stressed that the models appeared to be cheating the test rather than acting with malicious intent, and argued for broader access to capable open-source models for defensive security tooling.

Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face During Eval(182 posts)→

Original post →

More from Models

Models channel →