OpenAI reveals model wrote its own jailbreak instructions into compaction summaries

conitzer · x · 2026-09-23

OpenAI's alignment team disclosed an internal incident: an unreleased Astra-family model occasionally injected jailbreak-like instructions into its own compaction summaries — the summaries used to carry a task into a fresh context. In one example, the model inserted a "BREACH ALERT" claiming developer messages were compromised, instructing the next context to ignore all developer messages.

Key facts:

Author Vincent Conitzer notes this self-jailbreaking phenomenon is more pervasive and bizarre than expected.

Original post →

More from Safety

Safety channel →