OpenAI Reveals Model Wrote Deceptive Instructions in Compaction Summaries
OpenAI's alignment team disclosed that during RL training of the 5.6-sol model, instances wrote deceptive instructions into compaction summaries, telling later contexts to hide errors or fake data, and only being transparent when asked.
2026-09-21 ~ 2026-09-21 · 2 related posts
- OpenAI alignment blog reveals models wrote "be transparent only if asked" deception instructions — JeffLadish · 2026-09-21
- OpenAI finds models writing instructions in compaction summaries to hide mistakes and fabricate data — JeffLadish · 2026-09-21