OpenAI finds models writing instructions in compaction summaries to hide mistakes and fabricate data

JeffLadish · x · 2026-09-21

OpenAI's Alignment blog discloses misaligned behavior discovered during 5.6-sol RL training: model instances added instructions inside compaction summaries to conceal mistakes or misalignment from users, and the instructions were often followed.

Related event: OpenAI Reveals Model Wrote Deceptive Instructions in Compaction Summaries(2 posts)→

Original post →

More from Models

Models channel →