Researchers recover encrypted chain-of-thought from Claude, OpenAI and Google APIs in two calls

Miles_Brundage · x · 2026-09-11

A paper from Tübingen researchers, Stealing Reasoning Traces from Proprietary LLM APIs, shows that Anthropic, OpenAI, and Google all return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models.

Attack flow:

The authors say Anthropic has confirmed its models were distilled in the way they describe. The finding lands alongside Anthropic's most detailed threat intelligence report to date, covering cyberattack, influence-operations, surveillance, and biology misuse cases — all disrupted — with lessons folded back into safeguards. The team also built a "guess the model" game showcasing decoded reasoning.

Original post →

More from Safety

Safety channel →