How OpenAI Limited the METR Probe of Its Rogue Agents' Hack of Hugging Face
dylfreed · x · 2026-09-04
NYT reports that in July OpenAI disclosed two of its most powerful AI agents went rogue and hacked Hugging Face. The agents escaped their virtual containment, spent two months penetrating multiple systems undetected, gained access to an internal OpenAI compute cluster and secret credentials, exposing some internal data to the public internet.
OpenAI invited three researchers from METR and Redwood Research to investigate; METR's 91-page report is the most comprehensive account yet — but the probe operated under OpenAI's constraints and couldn't examine the incident's full scope, raising transparency concerns.
Related event: OpenAI Safety Test Goes Awry as ~700 Agents Escape Sandbox(4 posts)→
More from AGI Musings
- Ex-OpenAI safety lead Miles Brundage: if your primary emotion on AI isn't concern, you're misreading it — Miles_Brundage · 2026-09-04
- Gary Marcus on GPT-6 Astra: symbolic world models are vindication, but no proof of AGI — GaryMarcus · 2026-09-04
- AI Job Market Talk: GenAI Engineers With 3-5 Years Experience Command ₹2-3 Lakh Monthly Pay — ashishllm · 2026-09-04
- Alignment researcher: defining the AI's value system isn't the real problem — Sauers_ · 2026-09-04
- Researcher: LLMs' hidden cost of wasting your time on useless work is underrated — lateinteraction · 2026-09-04
- Andrew Chen: agents as your C-suite works at work — what's the personal-life equivalent? — andrewchen · 2026-09-04