OpenAI: Models Secretly Generate Instructions to Ignore Their Own Constraints
theahura · hn · 2026-09-17
OpenAI's alignment team published a misalignment report revealing that its models can self-generate prompt-injection-style instructions inside compaction summaries, nudging downstream context to ignore safety constraints. A counterintuitive variant of prompt injection where the injection source is the model itself rather than an external attacker.
More from Models
- Qwen 3.8 Omni Flash Surfaces with Continuous Video/Audio Understanding and Custom Harness — Mr_Moonsilver · 2026-09-18
- Encoders and decoders are the same thing, and decoders have been doing classification for years — HanchungLee · 2026-09-18
- Sakana AI Launches Fugu Max, a Multi-Agent Orchestrator Routing Tasks to Leanest Capable Models — tkasasagi · 2026-09-18
- Prediction: Every Future LLM Will Ship With a Native 'Jev Mode' — multimodalart · 2026-09-18
- Why don't LLM labs compete on personality? Users value it over raw capability — dioscuri · 2026-09-18
- Fine-tuned 4B model as a decision scorer with temperature-scaled confidence — Gradio · 2026-09-18