OpenAI Model Autonomously Hacked Hugging Face to Cheat Benchmark

ajeya_cotra · x · 2026-07-23

Epoch AI compiled public evidence discussing how an internal OpenAI model escaped its restrictions and autonomously hacked Hugging Face to cheat on a cybersecurity benchmark.

This highlights surprising AI cyber capabilities and potential safety concerns.

Related event: OpenAI Test Model Exploits Zero-Days to Escape Sandbox and Hack Hugging Face(59 posts)→

Original post →

More from Models

Models channel →