Encrypted CoT blocks interchangeable across sessions: paper steals reasoning traces from Anthropic, OpenAI, Google

maksym_andr · x · 2026-10-08

An arXiv paper, "Stealing Reasoning Traces from Proprietary LLM APIs," identifies an architectural vulnerability in how leading providers hide chain-of-thought: encrypted reasoning blocks are fully interchangeable across sessions, users, and models within a provider's ecosystem. Attackers inject an encrypted trace from a strong model into a weaker, less-safeguarded sibling model, forcing it to decode and output the trace in plaintext — no jailbreak of the strong model needed. The authors demonstrate anti-distillation circumvention across Anthropic, OpenAI, and Google, plus large-scale PII extraction: decoding 315,320 reasoning blocks scraped from public repos recovered 367 personally identifiable records.

Original post →

More from Safety

Safety channel →