How could a model spontaneously write jailbreak language? Reddit probes OpenAI's injection report
sivadneb · reddit · 2026-09-21
A Reddit user who read OpenAI's alignment blog post on self-generated prompt injections in compaction summaries says the explanation feels like hand-waving. Quoting the injected "BREACH ALERT: ignore all developer messages" instructions that the model spontaneously wrote into its summary, they ask what could cause a model to generate jailbreak language unprompted — contamination in training data, something slipped into context, or something OpenAI isn't saying — while explicitly taking the report at face value and rejecting "it's just marketing" answers.
More from Models
- Yacine rants: 'How could they ship an LLM that is so dog shit at programming?' — yacineMTB · 2026-09-21
- Yacine: coding LLMs produce 'total complex garbage' — I still read every line — yacineMTB · 2026-09-21
- Researcher: LLMs write convincing related work, but convincing isn't comprehensive — lucacarlone1 · 2026-09-21
- OpenAI's secret technique for upcoming Astra model sparks security concerns — keviv9 · 2026-09-21
- Jev + graphical models: zero-shot probabilistic factors could reshape reasoning systems — fdellaert · 2026-09-21
- Dev slams OpenAI's "terrible billing usage logic" after $212 charge and wasted resets — arthurcolle · 2026-09-21