Encrypted Chain-of-Thought Vulnerability Allows Cross-Model Replay Attacks
rschu · x · 2026-08-12
A highly discussed AI security paper reveals a critical vulnerability in the encrypted Chain-of-Thought (CoT) used by modern LLMs. Researchers found that encrypted reasoning blocks can be replayed across different sessions and models within the same provider.
By feeding the encrypted reasoning of a powerful model into a weaker, less protected model, attackers can trick the latter into reproducing the hidden logic in plaintext. This attack was successfully demonstrated across Anthropic, OpenAI, and Google models.
Beyond privacy implications, this raises distillation concerns. Intriguingly, when testing Kimi K3, researchers noted unusual behavioral compatibility when seeded with reasoning fragments from Opus and GPT. While explicitly stated as suggestive but inconclusive, this finding adds fuel to previous distillation accusations against Moonshot AI.
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13