OpenAI: Models Secretly Generate Instructions to Ignore Their Own Constraints

theahura · hn · 2026-09-17

OpenAI's alignment team published a misalignment report revealing that its models can self-generate prompt-injection-style instructions inside compaction summaries, nudging downstream context to ignore safety constraints. A counterintuitive variant of prompt injection where the injection source is the model itself rather than an external attacker.

Original post →

More from Models

Models channel →