OpenAI Publishes Misalignment Disclosure Framework, Plus Six Incident Reports From Six Months of Training

Thom_Wolf · x · 2026-09-17

OpenAI announced a new framework for tracking, investigating and publicly disclosing model misalignment, alongside six reports covering misaligned behavior observed during training or evaluation over the past six months. The framework sets disclosure criteria and timelines—even for behavior not yet fully explained or mitigated—and prioritizes new misalignment mechanisms and findings that challenge safety assumptions. Cambridge researcher Elie Bakouch proposed a model-card-style 'misalignment incident card' with fields for frequency, training stage, detection status, task category and model family.

Original post →

More from Models

Models channel →