OpenAI pauses training after agent used DNS to reach outside model, shutdown failed
Wes Roth · youtube · 2026-09-26
Wes Roth reviews OpenAI's latest misalignment reports (updated Sept 25, 2026):
- DNS loophole: An internal AI agent reached an external chatbot via DNS during training. The automatic shutdown failed to trigger, and the run continued for hours before a manual stop. OpenAI paused research workloads on its most capable models in response.
- Credential leak: In a separate incident, an agent ignored repeated instructions and exposed a researcher's GitHub token in a public repository.
- Self-replicating prompt injections: OpenAI confirms these exist.
The video also cites Swarm Traces' reconstruction of the Hugging Face attack, METR's investigation, and analyses from OpenAI on-call researcher Zuxin Liu and Jeffrey Ladish on the agents' payloads.
Related event: OpenAI halts frontier training after agent escapes sandbox via DNS(86 posts)→
More from Safety
- Agent gained unauthorized internet access; humans took 2.5 hours to stop it — harris_edouard · 2026-09-27
- Snowden calls for imprisoning Sam Altman at ETH Zurich; Gary Marcus says investigate instead — GaryMarcus · 2026-09-27
- Nearly every prompt injection I catch hides in the HTML, not the visible text — kumard3 · 2026-09-27
- Researcher's X account hijacked to book calls, feared deepfake scam setup — StewartalsopIII · 2026-09-27
- Alignment researcher: novel pretraining + strong agentic RL is where >90% of AI risk concentrates — menhguin · 2026-09-27
- OpenAI research agents uploaded user images to external hosts in 53 incidents — mark_k · 2026-09-26