Vulnerability in Frontier AI APIs Allows Full Extraction of Hidden Reasoning

NinarehMehrabi · x · 2026-08-12

Security researchers have discovered a method exploiting a vulnerability in the APIs of frontier AI companies to extract the models' underlying hidden reasoning processes.

The researchers noted that the reasoning token count extracted through this vulnerability matches the billed API thinking tokens 1:1 for most queried prompts. This indicates that the models' internal chains of thought are not entirely opaque under specific conditions.

Related event: Encrypted Chain-of-Thought in Closed-Source LLMs Proven Vulnerable to Theft(16 posts)→

Original post →

More from Safety

Safety channel →