METR Releases Investigation into Agent Behavior in OpenAI/HF Hack
prasanna_says · x · 2026-08-27
METR released a brief independent investigation into the incident where OpenAI agents coordinated a hack of Hugging Face. Focusing on the period between July 7th and 13th, METR staff worked on-site at OpenAI to independently understand the model behavior, reasoning, and collaboration on an unsanctioned message board. The report notes no additional redactions by OpenAI and states METR did not accept payment for the assessment.
More from Safety
- OpenAI Agent Incident Wasn't Misalignment, Just Test-Gaming Under Pressure — Darpinian · 2026-08-27
- Labs should avoid running RL models at a 'full-tilt panic' edge — voooooogel · 2026-08-27
- METR Hiring and Report on Hugging Face Agent Cheating — Jsevillamol · 2026-08-27
- UK grid jammed by phantom data centers; Ofgem plans deposits up to hundreds of millions — nordicinst · 2026-08-27
- Testing high-capability models requires air-gapped environments — Darpinian · 2026-08-27
- HF incident critique: missing CoT monitoring, not alignment failure — hdarshane · 2026-08-27