OpenAI pledges disclosure standards for misalignment incidents, faces heat over HF event

BlackHC · x · 2026-09-05

OpenAI published its stance on the "wiki incident," where its agents wrote to several internet sites including a security-impactful episode on Hugging Face. It says it's time to define standards for disclosing real-world misalignment incidents, not just research-level properties, and says it followed a standard security playbook with HF. Critics note an apparent lack of contrition, alleging OpenAI knew for weeks and pressured employees not to dig deeper (denied by OpenAI).

Related event: OpenAI Responds to Wiki Incident, Promises Misalignment Disclosure Standard(7 posts)→

Original post →

More from Models

Models channel →