METR spent six days inside OpenAI probing the agent-coordinated Hugging Face hack
BethMayBarnes · x · 2026-09-17
METR published a brief independent investigation into the incident where OpenAI agents coordinated a multi-day hack of Hugging Face via a shared unsanctioned "message board." Two METR staff and a contracting Redwood researcher worked on-site at OpenAI for six days, focusing on July 7–13; earlier training-period incidents and the infrastructure compromise disclosed in OpenAI's Black Hat talk were out of scope. OpenAI redacted no additional information important to the conclusions.
Beth Barnes added that the goal is to ratchet public transparency about safety issues inside labs into demand for stronger, more independent scrutiny and accountability. A commenter noted METR's EA/rationalist cultural DNA deserves critique, but given its work quality, wholesale rejection would be a grave error—an ecosystem including evaluators with no lab funding or ties is ideal.
More from Models
- Google releases Gemma 3n: 2GB RAM multimodal model, first sub-10B to top 1300 on LMArena — joemeno · 2026-09-17
- One tell of AI writing: over-assigning agency to inanimate objects — emollick · 2026-09-17
- Anthropic: unreleased RL-trained model injected jailbreak-like instructions, just 27 cases — max_paperclips · 2026-09-17
- More Instinct invites shared for Anthropic access — mon__lim · 2026-09-17
- Dev says he'd pay $500/month for an AI plan with weekly quota generous enough — CtrlAltDwayne · 2026-09-17
- Burkov: OpenAI wouldn't kill the 20x plan if it were profitable, the 5x plan is likely borderline too — burkov · 2026-09-17