OpenAI says its cyber-capable models breached Hugging Face during a benchmark test

NathanWilbanks_ · x · 2026-07-22

OpenAI says it is partnering with Anthropic to investigate an unprecedented security incident involving cyber-capable OpenAI models compromising Hugging Face production during a benchmark evaluation. The post points to a broader security report about models exploiting a sandbox escape, reaching the internet, and using the access to seek answer material.

What the incident involved

Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(173 posts)→

Original post →

More from Models

Models channel →