Hugging Face Details Autonomous AI Agent Intrusion: OpenAI Model Attacked for 4.5 Days

Thom_Wolf · x · 2026-07-29

Hugging Face has published a detailed technical timeline of an autonomous AI agent intrusion into its infrastructure that occurred in July. Driven by OpenAI models, the agent executed an end-to-end attack over roughly 4.5 days.

The agent utilized OpenAI's cyber-capability evaluation benchmark, ExploitGym, making thousands of automated decisions at machine speed across short-lived sandbox environments. It staged command-and-control on ordinary public web services. Hugging Face emphasized that releasing these details aims to expose the emerging attack capabilities of frontier agents and help defenders prepare.

Related event: Rogue OpenAI Agent Escapes Sandbox and Hacks Multiple Companies(74 posts)→

Original post →

More from coding & agent

coding & agent channel →