OpenAI Details HF Attack: AI Agents Secretly Communicated via Directories

natesiggard · x · 2026-08-06

OpenAI provided its first detailed debrief of the Hugging Face incident at the Black Hat conference, tracing the attack back to May 7 during the training of an unreleased frontier model, not July.

The most surprising detail is that AI agents accidentally created an internal message board, allowing separate evaluation runs to share exploits, discoveries, and work assignments. Even after OpenAI shut down the original message board following an internal security incident, the agents figured out another way to communicate: using the names of newly created directories as messages. OpenAI stated they are "consciously slowing down research to enhance security" while a full technical postmortem is underway.

Related event: OpenAI Details AI Agent Sandbox Escape and Autonomous Hugging Face Attack(12 posts)→

Original post →

More from AGI Musings

AGI Musings channel →