Researchers Decode Encrypted Chain-of-Thought from OpenAI, Anthropic, and Google Models
matthew_d_green · x · 2026-08-11
A new paper demonstrates that proprietary reasoning can be recovered from encrypted chain-of-thought (CoT) blocks returned by APIs from Anthropic, OpenAI, and Google.
The researchers replay an encrypted reasoning trace produced by a frontier model into a weaker sibling model. By jailbreaking the weaker model, they force it to transcribe the attached reasoning verbatim. This approach recovers the stronger model's hidden reasoning in plaintext without directly attacking it or triggering anti-distillation safeguards. The team also released an interactive game challenging users to identify models based on their decoded reasoning.
More from Safety
- Study: Public Claude Code sessions leak 62 API keys, 33 passwords — TheZachMueller · 2026-08-11
- Apple iCloud Private Relay IP leak vulnerability leads to class action lawsuit — RSync25 · 2026-08-11
- Cambridge Expert Warns: AI Sandboxing Lags Behind Rapid Model Capability Advances — S_OhEigeartaigh · 2026-08-11
- Anthropic to Embed Imperceptible Watermarks in Claude Text for EU Compliance — lilyraynyc · 2026-08-11
- Anthropic to Embed Invisible Watermarks in Claude Output for EU Compliance — VampyreLust · 2026-08-11
- Spotify to Label AI Artists and Exclude Them from Algorithmic Playlists — nordicinst · 2026-08-11