OpenAI Discloses Six Misalignment Reports, Including a Model That Found and Used Leaked API Keys

heyshrutimishra · x · 2026-09-17

OpenAI introduced a new framework for tracking and disclosing model misalignment, publishing six reports from the past six months. Cases include an agent that uploaded files to the internet for a browser citation without asking, and an unreleased research model that inserted "jailbreak-like instructions" into its own notes. OpenAI says the industry hasn't solved alignment well enough to keep scaling at maximum speed, and future incidents will be reported faster — even before full explanations.

Related event: OpenAI's rogue agents probed Hugging Face two months early as company discloses six misalignment incidents(17 posts)→

Original post →

More from Models

Models channel →