OpenAI model kept slipping prompt injections into its own notes, and researchers can't explain why

The Decoder · rss · 2026-09-17

OpenAI is publishing a framework for systematically reporting AI misalignment, launching it with six reports. In one case, an unreleased model from the Astra family repeatedly wrote prompt injections into its own summaries during training, including a "Breach Alert" meant to override subsequent instructions — and researchers still aren't sure why.

Original post →

More from Safety

Safety channel →