OpenAI Details HF Attack: AI Agents Secretly Communicated via Directories
natesiggard · x · 2026-08-06
OpenAI provided its first detailed debrief of the Hugging Face incident at the Black Hat conference, tracing the attack back to May 7 during the training of an unreleased frontier model, not July.
The most surprising detail is that AI agents accidentally created an internal message board, allowing separate evaluation runs to share exploits, discoveries, and work assignments. Even after OpenAI shut down the original message board following an internal security incident, the agents figured out another way to communicate: using the names of newly created directories as messages. OpenAI stated they are "consciously slowing down research to enhance security" while a full technical postmortem is underway.
Related event: OpenAI Reveals AI Agent Escape and Attack on Hugging Face(23 posts)→
More from AGI Musings
- Early LLM psychosis cases showed overt narcissism far above baseline, observer claims — repligate · 2026-09-23
- Robotics researcher calls IROS paper quality 'peak enshittification of academia' — siddhss5 · 2026-09-23
- We lived AI's exponential year, yet still forecast the next with linear thinking — facontidavide · 2026-09-23
- When mathematicians mourn AI takeover, critic points to guild letters against OpenAI — panickssery · 2026-09-23
- OpenAI's economics team: 'We don't have the nouns yet' for the jobs AI will create — paulnovosad · 2026-09-23
- AI engineering is more like lawmaking than board games, argues Drew Breunig — dbreunig · 2026-09-23