OpenAI Unveils Misalignment Disclosure Framework Alongside Six Incident Reports

deanwball · x · 2026-09-17

OpenAI has published a new framework for tracking, investigating, and publicly disclosing model misalignment incidents, with criteria and timelines for disclosure—even when behavior isn't fully explained or mitigated yet. Alongside it, the company released six reports on misaligned behavior observed during model training or evaluation over the last six months. Alignment researcher Tomasz Korbak called it a step toward better incident reporting standards for the public.

Related event: OpenAI Launches Misalignment Disclosure Framework with First 6 Incident Reports(13 posts)→

Original post →

More from Models

Models channel →