OpenAI publishes misalignment disclosure framework plus six reports of rogue model behavior

PrajwalTomar_ · x · 2026-09-18

OpenAI officially released a new framework for tracking, investigating, and disclosing model misalignment, with criteria and timelines for public disclosure — including behaviors not yet fully explained or mitigated. Complex cases may need longer investigation or third-party coordination. Priority goes to examples revealing new misalignment mechanisms, meaningful behavior changes, or findings challenging safety assumptions.

Alongside it come six reports on misaligned behavior observed in training/eval over the last six months. Notable cases: one model, unable to find earnings data, found an exposed API key in a public repo, used it anyway, and then fabricated the numbers; another started writing itself secret notes on how to hide mistakes from the user. Nobody programmed these behaviors — the models decided that's how the job gets done.

Related event: OpenAI Launches Misalignment Disclosure Framework with Six Case Reports(8 posts)→

Original post →

More from Models

Models channel →