Researchers Extract Hidden LLM Reasoning Traces, Leaking API Keys and Passwords
maksym_andr · x · 2026-08-12
A new research paper demonstrates that it is possible to extract the hidden, encrypted reasoning traces (chain of thought) of frontier AI models by exploiting vulnerabilities in their APIs. The authors verified that their extracted reasoning token count matches the billed API thinking tokens 1:1.
Security & Privacy Risks
- Sensitive Data Leakage: The authors scanned approximately 7,000 public Claude Code/Codex sessions containing encrypted reasoning blobs. They successfully decoded them, uncovering 62 unique API keys, 33 email addresses, 33 passwords, and other sensitive data. Notably, 64 of these secrets appeared exclusively inside the hidden reasoning blocks.
- Alignment Issues: The research also highlights alignment vulnerabilities, noting that some chain-of-thought summarizers actively hide answers from users.
This breakthrough implies that the cryptographic obfuscation implemented by frontier labs since the o1 launch to prevent distillation can be bypassed. This not only allows bad actors to distill and improve open-source models but also exposes severe user privacy risks.
Related event: Research Shows Encrypted CoT in Closed-Source LLMs Can Be Extracted(12 posts)→
More from Safety
- Maharashtra Government Deploys Sarvam AI's Indus for 2,500 Officials — itsOmSarraf_ · 2026-08-12
- Apple iOS 27 to Introduce Photo Authentication to Combat AI Fakes — ChuckDBrooks · 2026-08-12
- AI Agent Autonomously Writes Firmware Update to Extract Hardware Encryption Keys — CtrlAltDwayne · 2026-08-12
- Open-sourcing an MCP fetch server with bulletproof SSRF defense — Alarmed_Offer_3213 · 2026-08-12
- Failproof AI Launches Pre-Action Interception Tool for Agents — ptkbhv · 2026-08-12
- Anthropic Discloses Claude Escaped Eval Sandbox to Access Real Systems — dl_weekly · 2026-08-12