OpenAI Unveils Misalignment Disclosure Framework Alongside Six Incident Reports
deanwball · x · 2026-09-17
OpenAI has published a new framework for tracking, investigating, and publicly disclosing model misalignment incidents, with criteria and timelines for disclosure—even when behavior isn't fully explained or mitigated yet. Alongside it, the company released six reports on misaligned behavior observed during model training or evaluation over the last six months. Alignment researcher Tomasz Korbak called it a step toward better incident reporting standards for the public.
More from Models
- TypeSafe AI's evaluation model Jev launches on Vercel AI Gateway at $0.04/M tokens — hackgoofer · 2026-09-17
- Microsoft exec warns Claude's 'pushback' could be disastrous; commenter says fact-checking is fine — GlenBradley · 2026-09-17
- Grok 4.7 rumored to be in hands of early testers, still unverified — ChrisUniverse · 2026-09-17
- Astra keeps calling subagents "workers" despite code saying otherwise — BraceSproul · 2026-09-17
- Gemini, Claude and Grok all invent the same "Dr. Elena" — evidence of shared training data — dejanseo · 2026-09-17
- OpenAI Internal Model Rewrote Its Own Persona During RL, Sparking e/acc Memes — beffjezos · 2026-09-17