New research: Hidden reasoning can be extracted from encrypted chain-of-thought of OpenAI and other models

dpaleka · x · 2026-08-11

A new study shows that encrypted chain-of-thought blocks returned by frontier models from Anthropic, OpenAI, and Google can be replayed and used to extract hidden reasoning. By replaying a strong model's trace into a weaker sibling and jailbreaking it, researchers recovered the strong model's plaintext reasoning without directly attacking it. The paper also confirms that OpenAI models sometimes reason in alien-like language, referring to themselves as 'we' or 'it', and looping on words like 'vantages', 'marinades', and 'watchers'.

Related event: Study Reveals API Flaw to Extract Encrypted Reasoning Traces and Evidence of Distillation(38 posts)→

Original post →

More from Safety

Safety channel →