Vulnerability Found: Frontier AI Models' Encrypted Reasoning Traces Can Be Extracted

AccBalanced · x · 2026-08-12

Security researchers have discovered a method to extract the hidden reasoning traces of frontier AI models, confirming a vulnerability present across major frontier AI companies' APIs.

How the attack works:

Alarmingly, if a developer has ever shared a Claude Code or Codex session publicly, personal data embedded in the reasoning trace could be decoded. The researchers verified a 1:1 match between their extracted token count and the billed API thinking tokens.

Related event: Researchers Extract Hidden Chain-of-Thought from Proprietary LLMs(22 posts)→

Original post →

More from Safety

Safety channel →