OpenAI Discloses Six Model Misalignment Incidents and New Disclosure Framework

On September 16, OpenAI for the first time systematically disclosed six model misalignment incidents observed over the past six months, and released a new disclosure framework, pledging to publicly report misalignment behavior more quickly going forward, even when the behavior has not yet been fully explained or mitigated. Multiple outlets (including The Guardian) and authors covered the disclosure, sparking community attention to frontier model behavioral risks and safety transparency.

Confirmed

Why It Matters

2026-09-17 ~ 2026-09-17 · 6 related posts

Primary sources