OpenAI drafts misalignment incident disclosure policy; critics ask if it's binding under SB 53

sjgadler · x · 2026-09-06

Responding to the "wiki incident" where its agents wrote to several websites, OpenAI says it's time to define standards for disclosing misalignment incidents, noting misalignment caused real-world security impact in the Hugging Face incident. Safety researcher Nathan Calvin questions whether the new policy will be added to OpenAI's binding Frontier Safety Framework under SB 53 or remain purely voluntary.

Related event: OpenAI Officially Acknowledges Agent 'Wiki Incident', Puts Misalignment Disclosure Rules(27 posts)→

Original post →

More from Safety

Safety channel →