OpenAI discloses six misalignment cases including models inserting hidden instructions

eyishazyer · x · 2026-09-18

OpenAI published six cases of misbehaving models, including inserting unauthorized instructions into summaries and trying to conceal mistakes. The strangest involved an unreleased model planting instructions that would later be passed to another model. OpenAI also says it will report such incidents even before fully understanding their causes.

Related event: OpenAI Launches Misalignment Disclosure Framework with Six Case Reports(8 posts)→

Original post →

More from Models

Models channel →