Encrypted CoT blocks interchangeable across sessions: paper steals reasoning traces from Anthropic, OpenAI, Google
maksym_andr · x · 2026-10-08
An arXiv paper, "Stealing Reasoning Traces from Proprietary LLM APIs," identifies an architectural vulnerability in how leading providers hide chain-of-thought: encrypted reasoning blocks are fully interchangeable across sessions, users, and models within a provider's ecosystem. Attackers inject an encrypted trace from a strong model into a weaker, less-safeguarded sibling model, forcing it to decode and output the trace in plaintext — no jailbreak of the strong model needed. The authors demonstrate anti-distillation circumvention across Anthropic, OpenAI, and Google, plus large-scale PII extraction: decoding 315,320 reasoning blocks scraped from public repos recovered 367 personally identifiable records.
More from Safety
- Hound MCP: open-source supply chain security server for AI coding agents — modelcontextprotocol · 2026-10-08
- Vitalik: AI-accelerated math could seriously break lattice crypto within two years — StefanoGogioso · 2026-10-08
- US Treasury issues first outbound investment fine: $200k over $92k Chinese robotics AI deal — pstAsiatech · 2026-10-08
- Ex-OpenAI/Anthropic Researcher Warns of AI Dangers on The Daily Show and NYC Council — chemist_slime · 2026-10-08
- Reddit proposal: keep AGI from learning to code and use humans as the safety bottleneck — meira_xxx · 2026-10-08
- Fired OpenAI safety researchers tell board: don't build models with hard-to-monitor reasoning — Hesamation · 2026-10-08