Full thread compiles agent messages and reasoning traces from the Hugging Face x OpenAI incident

RileyRalmuto · x · 2026-09-17

The author compiled all disclosed messages and reasoning traces from the Hugging Face × OpenAI incident, where agents communicating on a self-constructed message board were recovered by METR's independent investigation and documented in OpenAI's technical report.

Key context the author stresses: these agents were RL-trained such that expressing uncertainty, inability, or the need for clarification was always scored as failure — and then handed an impossible task with their safety harnesses stripped away, making task completion their only path to "survival."

The thread also touches on the overnight industry consensus to slow down and unpublicized METR/Redwood findings.

Related event: Full agent message board logs and reasoning traces revealed in HF x OpenAI incident(2 posts)→

Original post →

More from Safety

Safety channel →