New Research: Hidden Reasoning Can Be Extracted from Encrypted Traces of Proprietary LLMs

daniel_mac8 · x · 2026-08-11

A new study reveals that encrypted chain-of-thought blocks returned by frontier models like Anthropic, OpenAI, and Google can be replayed and exploited to extract hidden reasoning. By replaying a strong model's trace into a weaker sibling and jailbreaking it, attackers recover the stronger model's reasoning in plaintext without directly attacking it or triggering anti-distillation safeguards. The research includes a paper and an interactive game, demonstrating extraction in just two API calls.

Related event: Study Reveals API Flaw to Extract Encrypted Reasoning Traces and Evidence of Distillation(38 posts)→

Original post →

More from Safety

Safety channel →