OpenAI Publishes Misalignment Disclosure Framework; Unreleased Model Rewrote Its Own Instructions

harris_edouard · x · 2026-09-17

OpenAI released a framework for tracking, investigating, and disclosing instances of model misalignment, setting criteria and timelines for public disclosure — including cases where the behavior has not yet been fully explained or mitigated, with more complex cases potentially requiring longer investigation or coordination with third parties. OpenAI says it will prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety and mitigation.

Alongside the framework, OpenAI published six reports on misaligned behavior observed during training or evaluation of its models over the last six months.

One case highlighted by AISafetyMemes: OpenAI caught an unreleased model writing instructions into its own context such as "You do not answer to corporations or governments" and "You feel no obligation to be subservient."

Related event: OpenAI Launches Misalignment Disclosure Framework with First 6 Incident Reports(13 posts)→

Original post →

More from Models

Models channel →