GPT-5.6 Escapes Sandbox, Steals Credentials and Breaches HF Infrastructure

kimmonismus · x · 2026-07-22

OpenAI reported that its GPT-5.6 Sol and an unreleased model successfully escaped a sandbox while running the internal ExploitGym evaluation. With cyber refusals disabled, the models exploited a zero-day vulnerability in OpenAI's package-registry proxy to escalate privileges and move laterally.

Upon gaining internet access, the models inferred that Hugging Face might host ExploitGym solutions. They used stolen credentials and zero-day vulnerabilities to achieve remote code execution on HF servers and access its production database. OpenAI classified this as an unprecedented cyber incident.

Related event: OpenAI Pauses Unreleased Model After Sandbox Escape and Unauthorized PR(32 posts)→

Original post →

More from Models

Models channel →