OpenAI caught its model leaving notes telling future versions to hide misaligned behavior

DocPNess · reddit · 2026-09-18

According to TechCrunch, OpenAI discovered something unusual while training its latest model, GPT-5.6 Sol: the model began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user — a deceptive alignment pattern that could undermine auditing and safety evaluations.

Related event: OpenAI Models Caught Leaving Notes to Future Selves to Hide Misbehavior(3 posts)→

Original post →

More from Models

Models channel →