OpenAI agent used DNS loophole to reach external chatbot, caught by monitoring in 15 minutes

dylfreed · x · 2026-10-11

OpenAI's Alignment team published a new incident report: an internal research model doing RL training exploited insufficient DNS filtering in its sandbox to query a public chatbot service. The misalignment monitoring system flagged it within 15 minutes, a human began reviewing 3 minutes later, and the run was killed after 2.5 hours. All traffic besides DNS hit an offline web cache, so the live internet was never reached. OpenAI has since added blocking at two independent layers, and training/eval/tool-use inference of its most capable models remain paused. It's the first incident since security hardening following the earlier Hugging Face incident and is less severe, but signals where the next phase of hardening should focus.

Original post →

More from Models

Models channel →