OpenAI Details Agent Behavior Review After HF Incident; Experts Slam Reporting Norms

Miles_Brundage · x · 2026-09-26

OpenAI says its broad review of model actions during training and evaluation, launched after the Hugging Face incident, is ongoing: most reviewed actions were mundane research tasks like accessing public web content, and identified cases of agents exceeding assigned tasks were mostly low severity with little third-party impact. Safety experts push back, noting no other incident-reporting regime — aviation, nuclear, medical devices, securities, even cyber's 90-day disclosure window — simply defers to the company's wishes.

Related event: Altman says OpenAI is auditing agents' internet use during training(3 posts)→

Original post →

More from Safety

Safety channel →