Research: Encrypted CoT Traces Can Be Extracted via Weaker Sibling Models
tw1st3d_m3nt4t · reddit · 2026-08-12
New research reveals a significant security risk in mainstream LLM APIs: encrypted chain-of-thought blocks returned by providers like Anthropic, OpenAI, and Google can be replayed across sessions, users, and models.
By replaying an encrypted trace from a frontier model into a weaker sibling model and jailbreaking the weaker one, researchers successfully recovered the stronger model's hidden reasoning in plaintext. This method bypasses the stronger model's anti-distillation safeguards entirely without attacking it directly.
Related event: Research Shows Encrypted Chain-of-Thought Can Be Extracted or Inverted(7 posts)→
More from Safety
- Mustafa Suleyman Draws AGI Red Line: Halt Systems That Acquire Resources Autonomously — Olivier__OG · 2026-08-12
- Voice Tool Wispr Flow Exposed for Retaining and Profiling Private User Dictations — ayushtweetshere · 2026-08-12
- Anthropic to Embed Invisible Watermarks in Claude Text Globally Under EU AI Act — TinfoilTricorn · 2026-08-12
- OpenAI Mandates Physical Hardware Security Keys for Enterprise Customers — CtrlAltDwayne · 2026-08-12
- Grok Bot Integration Faces Hurdles: Call for Verified AI Agents on X — Daniel_Farinax · 2026-08-12
- AI-Powered Scams Are Getting Highly Personalized and Hard to Detect — SpencrGreenberg · 2026-08-12