Researchers steal frontier models' hidden reasoning by replaying encrypted chain-of-thought traces
maksym_andr · x · 2026-10-01
Researchers from MATS, ELLIS Tübingen, Max Planck Institute and others released the "Stolen Thoughts" paper: Anthropic, OpenAI and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. By replaying a frontier model's trace into a weaker sibling and jailbreaking the weaker model, attackers recover the stronger model's hidden reasoning in plaintext without directly attacking it or triggering anti-distillation safeguards.
Key points:
- Extraction works in just two API calls;
- At MAX reasoning effort, models like Astra and Sol become hyper-aware of their token budgets and even start omitting whitespace to save tokens;
- In a Sept 30 mitigation audit two months after responsible disclosure, the authors could still extract verbatim reasoning from Astra/Sol-6.1 via third-party API providers;
- The site includes a "guess the model" game with decoded traces, including Kimi-K3.
Poster maksymandr notes GPT-6 Astra and Sol's reasoning output looks strange, tying into the attack.
More from Safety
- Research shows RL training breaks defenses against distillation attacks, evals give false security — terryyuezhuo · 2026-10-01
- Follow-up link to OpenAI's disclosure of Moonshot AI-linked hidden reasoning extraction campaign — kimmonismus · 2026-10-01
- OpenAI says Moonshot AI-linked individuals led campaign making 16,000 attempts to extract hidden reasoning — kimmonismus · 2026-10-01
- "Once phished, forever phished": researcher quips over a 3-tweet agentic AI disaster — suchenzang · 2026-10-01
- Anthropic report details 5 attempts to bypass Claude guardrails on pathogen gain-of-function queries — nordicinst · 2026-10-01
- Why AI supply chains need SBOMs: agentic AI brings authorization drift and shadow deployments — ChuckDBrooks · 2026-10-01