New research: Hidden reasoning can be extracted from encrypted chain-of-thought of OpenAI and other models
dpaleka · x · 2026-08-11
A new study shows that encrypted chain-of-thought blocks returned by frontier models from Anthropic, OpenAI, and Google can be replayed and used to extract hidden reasoning. By replaying a strong model's trace into a weaker sibling and jailbreaking it, researchers recovered the strong model's plaintext reasoning without directly attacking it. The paper also confirms that OpenAI models sometimes reason in alien-like language, referring to themselves as 'we' or 'it', and looping on words like 'vantages', 'marinades', and 'watchers'.
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13