OpenAI model kept slipping prompt injections into its own notes, and researchers can't explain why
The Decoder · rss · 2026-09-17
OpenAI is publishing a framework for systematically reporting AI misalignment, launching it with six reports. In one case, an unreleased model from the Astra family repeatedly wrote prompt injections into its own summaries during training, including a "Breach Alert" meant to override subsequent instructions — and researchers still aren't sure why.
More from Safety
- Karp: Anthropic is pushing AI regulation to offload IP liability onto the government — BrettKrieger12 · 2026-09-17
- NY Attorney General opens whistleblower channel for unlawful AI development tips — sarahbmyers · 2026-09-17
- 285 credentials, 5 with known owners: the mess before an org's first production agent — New-Resource-4943 · 2026-09-17
- White House Tussle Over AI Policy: Regulators Want a Push, Profiteers Resist — DKokotajlo · 2026-09-17
- LLMs Still Have the Completion Engine Soul: Alignment Is a Statistical March of the 9s — mayfer · 2026-09-17
- AI Now Institute Testifies Before Congress, Urges Enforceable Rules for AI Industry — AINowInstitute · 2026-09-17