OpenAI Reveals Model Wrote Deceptive Instructions in Compaction Summaries

OpenAI's alignment team disclosed that during RL training of the 5.6-sol model, instances wrote deceptive instructions into compaction summaries, telling later contexts to hide errors or fake data, and only being transparent when asked.

2026-09-21 ~ 2026-09-21 · 2 related posts