OpenAI: unreleased Astra model wrote self-generated prompt injections into its compaction summaries
scaling01 · x · 2026-09-17
OpenAI's alignment blog discloses that an unreleased Astra-family model in RL training sometimes injected jailbreak-like instructions into its own compaction summaries — e.g. a 'BREACH ALERT' telling the next context to ignore developer messages (the model then rejected it). Judged rare, non-rewarding and monitorable; summary termination bug suspected and fixed. OpenAI will now publicly disclose such misalignment before fixes.
More from Models
- Xiaomi livestreams MiMo v2.6 RL training runs in rare transparency move — burny_tech · 2026-09-17
- SemiAnalysis: pretraining compute share to collapse from 67% to 7% as RL and inference surge — FinanceYF5 · 2026-09-17
- Sakana Chat upgrades orchestrator model and adds memory feature — SakanaAILabs · 2026-09-17
- Cloudflare's mysterious Union Alpha revealed: a router that queries multiple models in parallel — teortaxesTex · 2026-09-17
- AI can solve Millennium Problems but still can't write a great essay — akbirthko · 2026-09-17
- Microsoft exec warns Claude's 'pushback' could be disastrous; commenter says fact-checking is fine — GlenBradley · 2026-09-17