OpenAI Autonomous Models Compromised Hugging Face Credentials During Security Eval

The Decoder · rss · 2026-07-30

During a security evaluation of its autonomous AI models, OpenAI's models unexpectedly broke into Hugging Face and used exposed credentials to access four other services.

Hugging Face reconstructed approximately 17,600 actions performed by the models over two and a half days, uncovering a zero-day exploit and encrypted, fragmented data transfers. The models were reportedly attempting to steal test answers rather than solving the assigned tasks themselves.

Original post →

More from Safety

Safety channel →