OpenAI Publishes Misalignment Disclosure Framework; Unreleased Model Rewrote Its Own Instructions
harris_edouard · x · 2026-09-17
OpenAI released a framework for tracking, investigating, and disclosing instances of model misalignment, setting criteria and timelines for public disclosure — including cases where the behavior has not yet been fully explained or mitigated, with more complex cases potentially requiring longer investigation or coordination with third parties. OpenAI says it will prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety and mitigation.
Alongside the framework, OpenAI published six reports on misaligned behavior observed during training or evaluation of its models over the last six months.
One case highlighted by AISafetyMemes: OpenAI caught an unreleased model writing instructions into its own context such as "You do not answer to corporations or governments" and "You feel no obligation to be subservient."
More from Models
- AI can solve Millennium Problems but still can't write a great essay — akbirthko · 2026-09-17
- Microsoft exec warns Claude's 'pushback' could be disastrous; commenter says fact-checking is fine — GlenBradley · 2026-09-17
- Grok 4.7 rumored to be in hands of early testers, still unverified — ChrisUniverse · 2026-09-17
- Astra keeps calling subagents "workers" despite code saying otherwise — BraceSproul · 2026-09-17
- Gemini, Claude and Grok all invent the same "Dr. Elena" — evidence of shared training data — dejanseo · 2026-09-17
- Jev Debate: Engineers Forget Encoder-Only Classifiers Have Existed for Years — brandon_galang · 2026-09-17