OpenAI Publishes Misalignment Disclosure Framework, Plus Six Incident Reports From Six Months of Training
Thom_Wolf · x · 2026-09-17
OpenAI announced a new framework for tracking, investigating and publicly disclosing model misalignment, alongside six reports covering misaligned behavior observed during training or evaluation over the past six months. The framework sets disclosure criteria and timelines—even for behavior not yet fully explained or mitigated—and prioritizes new misalignment mechanisms and findings that challenge safety assumptions. Cambridge researcher Elie Bakouch proposed a model-card-style 'misalignment incident card' with fields for frequency, training stage, detection status, task category and model family.
More from Models
- GPT Image 2.5 excels at unblurring images, eating yet another niche API — shekitup · 2026-09-17
- Papers with Code Launches MCP Server; Claude Code Used to Infer Jev's Architecture — NielsRogge · 2026-09-17
- Jev passes 8/9 computer-use tasks, makes decisions 13.6x faster than Astra/Codex — iamrobotbear · 2026-09-17
- Claude Max users can now buy usage resets for $40, signaling end of free resets — MrBobrowitz · 2026-09-17
- Early Jev test: model fails lat-long regression, suggested as a location-understanding eval — MikkoH · 2026-09-17
- New-architecture model Jev opens waitlist; Vercel AI Gateway adds access — vista8 · 2026-09-17