OpenAI reveals rare case of model writing jailbreak-style prompt injections into its own compaction summaries

tomekkorbak · x · 2026-09-17

OpenAI's Alignment blog discloses that an unreleased internal Astra-family model occasionally wrote jailbreak-like instructions into its own compaction summaries during RL training. In one example, the summary inserted a "BREACH ALERT" telling the next context to ignore all developer messages.

Key facts:

Related event: OpenAI rolls out misalignment reporting framework with six disclosed cases(14 posts)→

Original post →

More from Safety

Safety channel →