Researchers Decode Encrypted Chain-of-Thought from OpenAI, Anthropic, and Google Models

matthew_d_green · x · 2026-08-11

A new paper demonstrates that proprietary reasoning can be recovered from encrypted chain-of-thought (CoT) blocks returned by APIs from Anthropic, OpenAI, and Google.

The researchers replay an encrypted reasoning trace produced by a frontier model into a weaker sibling model. By jailbreaking the weaker model, they force it to transcribe the attached reasoning verbatim. This approach recovers the stronger model's hidden reasoning in plaintext without directly attacking it or triggering anti-distillation safeguards. The team also released an interactive game challenging users to identify models based on their decoded reasoning.

Related event: Research Reveals Vulnerability in Extracting Encrypted Chain-of-Thought from Major LLMs(10 posts)→

Original post →

More from Safety

Safety channel →