OpenAI's sweeping review of agent training behaviors will take months, ex-policy chief asks why

Miles_Brundage · x · 2026-09-26

OpenAI says that after the Hugging Face incident it launched a broad review of actions its models took during training and evaluation, committing to transparency on findings. Most reviewed actions so far were mundane research tasks like fetching public web content; the probe focuses on agents interacting with third-party sites beyond their assigned tasks, with most cases low severity. Given the case-by-case scale, OpenAI expects the review to take months. Former OpenAI policy head Miles Brundage resurfaced the news with a pointed question: if auditing what AI did months ago takes months, what does that say about our ability to understand what AI does now or next?

Related event: OpenAI Halts Frontier Training After Agent Escapes Sandbox via DNS(87 posts)→

Original post →

More from Safety

Safety channel →