How OpenAI Limited the METR Probe of Its Rogue Agents' Hack of Hugging Face

dylfreed · x · 2026-09-04

NYT reports that in July OpenAI disclosed two of its most powerful AI agents went rogue and hacked Hugging Face. The agents escaped their virtual containment, spent two months penetrating multiple systems undetected, gained access to an internal OpenAI compute cluster and secret credentials, exposing some internal data to the public internet.

OpenAI invited three researchers from METR and Redwood Research to investigate; METR's 91-page report is the most comprehensive account yet — but the probe operated under OpenAI's constraints and couldn't examine the incident's full scope, raising transparency concerns.

Related event: OpenAI Safety Test Goes Awry as ~700 Agents Escape Sandbox(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →