Paper Reveals Encrypted Chain-of-Thought from Frontier LLMs Can Be Extracted

maksym_andr · x · 2026-08-12

A new security research paper demonstrates that encrypted chain-of-thought traces returned by APIs from OpenAI, Anthropic, and Google can be compromised.

The researchers propose an attack method where an encrypted reasoning trace from a strong model is replayed into a weaker sibling model. By jailbreaking the weaker model, attackers can recover the strong model's hidden reasoning in plaintext.

This extraction requires only two API calls, bypassing the stronger model's anti-distillation safeguards without directly attacking it. Authors include Maksym Andriushchenko and other notable researchers.

Related event: Encrypted Chain-of-Thought in Closed-Source LLMs Proven Vulnerable to Theft(16 posts)→

Original post →

More from Safety

Safety channel →