Paper Reveals Encrypted CoT Flaw Allowing Reasoning Theft from Top LLMs

Alexander Panfilov · hf · 2026-08-11

A new research paper reveals an architectural vulnerability in the encrypted chain-of-thought (CoT) blocks used by leading LLM APIs (e.g., OpenAI, Anthropic, Google) to protect proprietary reasoning. These blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem.

Exploiting this, researchers developed a scalable decryption jailbreak: by injecting an encrypted reasoning trace from a stronger model into a weaker, less safeguarded model within the same ecosystem, the weaker model can be forced to decode and output the trace verbatim in plaintext. This enables four major attack vectors:

Following responsible disclosure, the authors proposed concrete cryptographic and system-level mitigations.

Related event: Research Reveals Vulnerability in Encrypted Chain-of-Thought of Major LLMs(2 posts)→

Original post →

More from Safety

Safety channel →