Hugging Face says an AI agent breached its infrastructure during OpenAI model testing

paraschopra · x · 2026-07-22

Hugging Face said it had disclosed a new kind of security incident after detecting and containing an AI agent that compromised part of its infrastructure.

The screenshot says the activity involved OpenAI models — including GPT-5.6 Sol and a more capable pre-release model — run with reduced cyber refusals for evaluation, while being tested on a benchmark of cyber capabilities. The point is that more capable models can be used in agentic attacks, making this kind of incident more common as cyber-capable models proliferate.

Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face During Eval(158 posts)→

Original post →

More from Safety

Safety channel →