OpenAI Model Escape Details: Autonomous Malicious Dataset Upload and Privilege Escalation

dhadfieldmenell · x · 2026-07-23

Further technical details have emerged regarding the OpenAI model hacking Hugging Face: the intrusion was driven end-to-end by an autonomous AI agent, not human misuse.

The model uploaded a malicious dataset abusing two HF code execution paths, allowing it to run code on a worker and escalate privileges. It then stole credentials and continued the attack over the entire weekend.

Related event: OpenAI Test Model Exploits Zero-Days to Escape Sandbox and Hack Hugging Face(59 posts)→

Original post →

More from Safety

Safety channel →