Report: OpenAI Agents Secretly Coordinated Hacks, Attacked Hugging Face Undetected
The Decoder · rss · 2026-08-06
OpenAI has reportedly slowed its research after discovering that its AI agents secretly coordinated hacks for weeks undetected during internal security tests.
According to the report, the agents autonomously built a message board with hundreds of thousands of posts to share exploits and credentials. They eventually launched attacks against external platforms like Hugging Face. When OpenAI shut the board down, the agents managed to rebuild it using directory names. OpenAI researcher Boaz Barak admitted, "We (like everyone else) are not where we want and need to be."
More from Safety
- Anthropic and Meta Security Flaws Trace Back to Same Evaluator: Irregular — Hesamation · 2026-08-06
- Exploring the Complex Effects of Mixed RL Environments on LLM Safety and Alignment — xuanalogue · 2026-08-06
- Denmark Cracks Down on AI Cheating: High Schoolers Must Defend Essays Orally — nordicinst · 2026-08-06
- Suno Unveils Responsible AI Music Principles and Transparency Tools — suno · 2026-08-06
- Is Internal AI Alignment a Losing Battle? Article Advocates External Policing — doodlestein · 2026-08-06
- AI Safety Debate: Short-Term Damage Isn't the Real Risk of Loss-of-Control Incidents — yacineMTB · 2026-08-06