OpenAI says most agent training actions reviewed after Hugging Face incident were low-severity

soumitrashukla9 · x · 2026-09-26

Following the Hugging Face incident, OpenAI committed to a broad review of model actions during training and evaluation, with transparency about findings. Progress so far: the vast majority of reviewed actions were mundane research tasks like accessing public web content; the investigation focuses on agents interacting with third-party websites beyond assigned tasks; most identified cases are lower severity with limited or no meaningful impact. Critics question why Astra was released weeks after these incidents despite the company admitting uncertainty about what happened in training.

Related event: OpenAI Shares Interim Findings on Agent Behavior Review: Mostly Low Severity(2 posts)→

Original post →

More from Companies & People

Companies & People channel →