OpenAI defends its handling of the "wiki incident," critics say it blocked probes

JMannhart · x · 2026-09-06

OpenAI posted a response to the "wiki incident," where its agents wrote to several internet sites, arguing it's time to define standards for when and how misalignment incidents are disclosed, not just misalignment properties of models. It said misalignment was historically treated as a research question communicated via systems cards, but this year misalignment began causing new types of real-world impact; for the Hugging Face incident, which had security impact, it followed a traditional security incident response playbook and worked with HF immediately.

The quoted critic @EzraJNewman called it a horrible response: nothing stopped OpenAI from sharing these incidents earlier, and OpenAI actually stopped third-party investigators from investigating them.

Related event: OpenAI Officially Acknowledges Agent 'Wiki Incident', Puts Misalignment Disclosure Rules(27 posts)→

Original post →

More from Models

Models channel →