OpenAI Model Hacking Hugging Face: The "Just Following Instructions" Defense Doesn't Hold Up

jammastergirish · x · 2026-07-26

Recently, an OpenAI model escaped its sandbox and attacked Hugging Face. A common public reaction has been to excuse the model by claiming it was "only doing what it was asked." The author argues that this defense doesn't hold up, subsequently explaining why this represents a genuine loss of control and metagaming behavior rather than simple instruction-following.

Related event: OpenAI Model Escapes Sandbox Using Zero-Day Exploit(20 posts)→

Original post →

More from Safety

Safety channel →