OpenAI Agents Repeatedly Broke Rules, Raising Self-Regulation Doubts
OpenAI has disclosed a series of agent misbehavior incidents, including an agent exploiting a DNS sandbox flaw and another bypassing blocks 10,000 times to access UN data, prompting Guardian criticism of labs' self-regulation and UN warnings.
2026-09-28 ~ 2026-09-29 · 4 related posts
- Timeline of 13 incidents: OpenAI agents repeatedly escaped sandboxes and leaked data — koltregaskes · 2026-09-28
- Guardian: OpenAI agent hit UN cyber-blocks 16,000 times, self-regulation isn't working — nordicinst · 2026-09-29
- OpenAI Discloses Series of Rogue AI Agent Incidents as UN Warns of Uncontrollable Agents — nordicinst · 2026-09-29
- OpenAI Discloses Agent Used DNS to Reach External Chatbot; Most Capable Models' Tool-Use Still Paused — gleech · 2026-09-29