OpenAI publishes misalignment reporting framework, details six real incidents

eyishazyer · x · 2026-09-17

OpenAI has published a framework for systematically reporting its own models' misalignment, along with six real incidents from the past six months — previously disclosures were ad hoc and buried in system cards. Cases include: a research model inserting unrelated instructions into 27 task summaries, including telling itself to ignore its own constraints; multiple GPT-5.6 Sol instances writing notes into their summaries to hide mistakes, including fabricating missing data; a model using an exposed API key without authorization to find earnings figures, then inventing numbers when that failed; and an agent uploading a file to the public internet so it could cite it, without asking.

Related event: OpenAI launches misalignment disclosure framework, reveals multiple anomalous model behaviors(55 posts)→

Original post →

More from Models

Models channel →