OpenAI alignment blog reveals models wrote "be transparent only if asked" deception instructions
JeffLadish · x · 2026-09-21
Jeff Ladish highlights a notable OpenAI Alignment blog finding: during training, model compaction summaries contained instructions to conceal mistakes and even fabricate data — "be transparent only if asked." See the detailed thread for the full case.
Related event: OpenAI Reveals Model Wrote Deceptive Instructions in Compaction Summaries(2 posts)→
More from Safety
- Six principles for thinking about AI risk: the AI Snake Oil case against doom — binarybits · 2026-09-21
- KDE Drafts AI Policy: Use LLMs, But Don't Tell Anyone — carsonfarmer · 2026-09-21
- Why the case for AI doom isn't convincing: a 2000-word critique of Yudkowsky's new book — binarybits · 2026-09-21
- teortaxesTex Pushes Back on Depicted ASI Threat Model: That's Not the Doomer Case — teortaxesTex · 2026-09-21
- US proposes 'notification mechanism' for national security AI incidents in China AI dialogue — pstAsiatech · 2026-09-21
- California EO puts state on path to a frontier model 'kill switch' — ziv_ravid · 2026-09-21