Researchers recover encrypted chain-of-thought from Claude, OpenAI and Google APIs in two calls
Miles_Brundage · x · 2026-09-11
A paper from Tübingen researchers, Stealing Reasoning Traces from Proprietary LLM APIs, shows that Anthropic, OpenAI, and Google all return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models.
Attack flow:
- Take an encrypted, signed thinking trace from a frontier model (e.g., Claude Opus);
- Replay it into a weaker sibling (e.g., Claude Haiku) with a jailbreak prompt asking the model to transcribe the attached reasoning verbatim;
- Recover the stronger model's hidden reasoning in plaintext without attacking it directly or tripping its anti-distillation safeguards.
The authors say Anthropic has confirmed its models were distilled in the way they describe. The finding lands alongside Anthropic's most detailed threat intelligence report to date, covering cyberattack, influence-operations, surveillance, and biology misuse cases — all disrupted — with lessons folded back into safeguards. The team also built a "guess the model" game showcasing decoded reasoning.
More from Safety
- Nearly 10% of exposed LiteLLM gateways accept default admin key 'sk-1234' — Thionne_WTZ · 2026-09-11
- Kelsey Piper: Labs plan to automate AI R&D with AI, shrinking human oversight within two years — round · 2026-09-11
- AI Evaluator Forum brings together Transluce, METR, RAND for independent AI evaluations — typewriters · 2026-09-11
- Anthropic says it stopped attempts to use models for potential biological weapons — connoraxiotes · 2026-09-11
- zetalyrae: extensional definitions of alignment only work retrospectively — zetalyrae · 2026-09-11
- "But China" is a legitimate concern in AI pacing debates, says Wildeford — peterwildeford · 2026-09-11