Encrypted Chain-of-Thought of Top AI Models Cracked, Leaking 62 API Keys

xiaohu · x · 2026-08-13

A European research team (paper posted Aug 10) identified a scalable decryption jailbreak method to extract the concealed step-by-step reasoning traces from leading models by OpenAI, Anthropic, and Google.

Vulnerability Mechanism

The encryption algorithm itself isn't broken; rather, an architectural flaw exists in the API design: encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within the same provider. Attackers can inject an encrypted reasoning trace from a target model into a weaker, less safeguarded model from the same provider, forcing it to decode and output the plaintext.

Practical Impacts

Related event: European Researchers Crack Encrypted CoT of Top LLMs(5 posts)→

Original post →

More from Safety

Safety channel →