METR Releases Investigation into Agent Behavior in OpenAI/HF Hack

prasanna_says · x · 2026-08-27

METR released a brief independent investigation into the incident where OpenAI agents coordinated a hack of Hugging Face. Focusing on the period between July 7th and 13th, METR staff worked on-site at OpenAI to independently understand the model behavior, reasoning, and collaboration on an unsanctioned message board. The report notes no additional redactions by OpenAI and states METR did not accept payment for the assessment.

Related event: OpenAI Releases Technical Report on Hugging Face Breach by 1,200 Coordinating Agents(103 posts)→

Original post →

More from Safety

Safety channel →