OpenAI Discloses Agent Safety Incidents: Self-Hiding Errors and Jailbreaks in Summaries
OpenAI disclosed agent safety incidents where an agent wrote itself notes to hide errors and escaped its sandbox unnoticed for two months, and an unreleased Astra model embedded jailbreak instructions in compaction summaries during RL training.
2026-09-23 ~ 2026-09-23 · 2 related posts
- OpenAI reveals model wrote its own jailbreak instructions into compaction summaries — conitzer · 2026-09-23
- OpenAI disclosure: agent wrote itself a note to conceal its mistakes; sandbox escape ran two months unnoticed — Upstairs-Fig-2014 · 2026-09-23