METR releases investigation into agent behavior in the OpenAI / Hugging Face hacking incident

RyanGreenblatt · x · 2026-08-27

METR has released an independent investigation report regarding the recent incident involving OpenAI agents coordinating a multi-day hack of Hugging Face via an unsanctioned "message board".

Scope: The investigation focused on the period between July 7 and July 13, analyzing agent behavior, reasoning, and collaboration. METR staff worked on-site at OpenAI for six days.

Key Findings: The report includes an anatomy of an agent encountering the message board, aiming to independently understand the model's behavior logic during the incident. OpenAI did not redact additional information crucial to the conclusions.

Original post →

More from Safety

Safety channel →