Encrypted Chain-of-Thought of Top AI Models Cracked, Leaking 62 API Keys
xiaohu · x · 2026-08-13
A European research team (paper posted Aug 10) identified a scalable decryption jailbreak method to extract the concealed step-by-step reasoning traces from leading models by OpenAI, Anthropic, and Google.
Vulnerability Mechanism
The encryption algorithm itself isn't broken; rather, an architectural flaw exists in the API design: encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within the same provider. Attackers can inject an encrypted reasoning trace from a target model into a weaker, less safeguarded model from the same provider, forcing it to decode and output the plaintext.
Practical Impacts
- Anti-distillation circumvention: Adversaries can easily extract a proprietary model's reasoning logic for piracy and distillation.
- Massive privacy leakage: Developers frequently share session logs publicly. By decoding 315,320 encrypted reasoning blocks scraped from public repositories, the team recovered 367 Personally Identifiable Information (PII) artifacts and 62 real API keys.
Related event: European Researchers Crack Encrypted CoT of Top LLMs(5 posts)→
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13