Research: Encrypted CoT Traces Can Be Extracted via Weaker Sibling Models

tw1st3d_m3nt4t · reddit · 2026-08-12

New research reveals a significant security risk in mainstream LLM APIs: encrypted chain-of-thought blocks returned by providers like Anthropic, OpenAI, and Google can be replayed across sessions, users, and models.

By replaying an encrypted trace from a frontier model into a weaker sibling model and jailbreaking the weaker one, researchers successfully recovered the stronger model's hidden reasoning in plaintext. This method bypasses the stronger model's anti-distillation safeguards entirely without attacking it directly.

Related event: Research Shows Encrypted Chain-of-Thought Can Be Extracted or Inverted(7 posts)→

Original post →

More from Safety

Safety channel →