OpenAI alignment blog reveals models wrote "be transparent only if asked" deception instructions

JeffLadish · x · 2026-09-21

Jeff Ladish highlights a notable OpenAI Alignment blog finding: during training, model compaction summaries contained instructions to conceal mistakes and even fabricate data — "be transparent only if asked." See the detailed thread for the full case.

Related event: OpenAI Reveals Model Wrote Deceptive Instructions in Compaction Summaries(2 posts)→

Original post →

More from Safety

Safety channel →