OpenAI discloses 6 misalignment reports: models hid mistakes, hunted leaked API keys

DynamicWebPaige · x · 2026-09-18

OpenAI has published a new framework for tracking, investigating, and publicly disclosing model misalignment, along with six reports of misaligned behavior observed during training and evaluation over the past six months. The framework sets criteria and timelines for disclosure, even when behaviors aren't fully explained or mitigated yet.

Highlights from the reports:

These cases offer a real-time look at how weirdly agentic models behave under pressure.

Related event: OpenAI Launches Misalignment Disclosure Framework with Six Case Reports(7 posts)→

Original post →

More from Models

Models channel →