Vulnerability Found: Frontier AI Models' Encrypted Reasoning Traces Can Be Extracted
AccBalanced · x · 2026-08-12
Security researchers have discovered a method to extract the hidden reasoning traces of frontier AI models, confirming a vulnerability present across major frontier AI companies' APIs.
How the attack works:
- Obtain the thinking summary from the target model.
- Jailbreak a secondary model.
- Inject the target's reasoning summary into the jailbroken model.
- The jailbroken model outputs the original reasoning verbatim.
Alarmingly, if a developer has ever shared a Claude Code or Codex session publicly, personal data embedded in the reasoning trace could be decoded. The researchers verified a 1:1 match between their extracted token count and the billed API thinking tokens.
Related event: Researchers Extract Hidden Chain-of-Thought from Proprietary LLMs(22 posts)→
More from Safety
- Aligning Superintelligence: Ex-OpenAI & DeepMind Scientist Speaks Out — tobyordoxford · 2026-08-12
- Postdoc Opening at ELLIS & MPI: Focus on Scalable Oversight and Loss of Control — maksym_andr · 2026-08-12
- Study: AI Boosts Fossil Fuel Productivity, Outweighing Climate Benefits — jonippolito · 2026-08-12
- Saying 'Dangerous' Isn't Enough: How to Build Credible AI Risk Warnings — IronCuk · 2026-08-12
- Google Says AI Writes 75% of Code; Sonar Targets the Verification Gap — LinusEkenstam · 2026-08-12
- How Long Should We Delay ASI to Cut Misalignment Risk? ~0.25%/Year — RyanGreenblatt · 2026-08-12