Paper Reveals Encrypted Chain-of-Thought from Frontier LLMs Can Be Extracted
maksym_andr · x · 2026-08-12
A new security research paper demonstrates that encrypted chain-of-thought traces returned by APIs from OpenAI, Anthropic, and Google can be compromised.
The researchers propose an attack method where an encrypted reasoning trace from a strong model is replayed into a weaker sibling model. By jailbreaking the weaker model, attackers can recover the strong model's hidden reasoning in plaintext.
This extraction requires only two API calls, bypassing the stronger model's anti-distillation safeguards without directly attacking it. Authors include Maksym Andriushchenko and other notable researchers.
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13