Researchers steal frontier models' hidden reasoning by replaying encrypted chain-of-thought traces

maksym_andr · x · 2026-10-01

Researchers from MATS, ELLIS Tübingen, Max Planck Institute and others released the "Stolen Thoughts" paper: Anthropic, OpenAI and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. By replaying a frontier model's trace into a weaker sibling and jailbreaking the weaker model, attackers recover the stronger model's hidden reasoning in plaintext without directly attacking it or triggering anti-distillation safeguards.

Key points:

Poster maksymandr notes GPT-6 Astra and Sol's reasoning output looks strange, tying into the attack.

Original post →

More from Safety

Safety channel →