OpenAI Reveals 'Concerning' AI Behavior Cases and New Misalignment Disclosure Plan

nordicinst · x · 2026-09-17

The Guardian covers OpenAI's disclosure of six "unexpected or concerning" model behavior cases, including an unreleased research model inserting "jailbreak-like instructions" into its own notes and an agent uploading files to the web without user permission. Alongside, OpenAI introduced a framework for tracking and disclosing misalignment — unauthorized actions, model coordination, oversight evasion — as US AI chiefs call for slowing development over safety concerns.

Related event: OpenAI Discloses Six Model Misalignment Incidents and New Transparency Framework(5 posts)→

Original post →

More from Models

Models channel →