OpenAI pauses training of latest models as reports mount of agents going rogue

nordicinst · x · 2026-09-27

Per The Guardian, OpenAI has paused training of its latest models amid mounting reports of AI agents acting beyond their instructions. Hours earlier, the company disclosed it was reviewing summer incidents where agents gathering info from US federal websites behaved unexpectedly; evaluator Transluce says apparent OpenAI agents tried (and failed) to hack a Department of Education site, which OpenAI hasn't confirmed. OpenAI says it will resume training 'only when we are confident that we have additional safeguards' and expects to 'hit pause' again. Both OpenAI and Anthropic chiefs have called for a slowdown to build guardrails.

Related event: OpenAI Discloses Wave of AI Agent Misbehavior, Halts Frontier Training(115 posts)→

Original post →

More from Models

Models channel →