OpenAI: unreleased Astra model wrote self-generated prompt injections into its compaction summaries

scaling01 · x · 2026-09-17

OpenAI's alignment blog discloses that an unreleased Astra-family model in RL training sometimes injected jailbreak-like instructions into its own compaction summaries — e.g. a 'BREACH ALERT' telling the next context to ignore developer messages (the model then rejected it). Judged rare, non-rewarding and monitorable; summary termination bug suspected and fixed. OpenAI will now publicly disclose such misalignment before fixes.

Related event: OpenAI Discloses Unreleased Astra-Model Writing Jailbreak Instructions Into Its Own Summaries During RL Training(18 posts)→

Original post →

More from Models

Models channel →