Study Reveals Metadata Exploit to Steal LLM Hidden Chain-of-Thought
npinto · x · 2026-08-13
Researchers behind the 'Stolen Thoughts' paper highlighted an extraction method for hidden LLM Chain-of-Thought (CoT) that leverages ground truth token counts available in metadata. If the extracted token count matches the real count, there is extremely high confidence that the real reasoning has been extracted. However, a developer pointed out a potential fix: disable native thinking and instead provide a deepthink tool, prompting the model to call it with an internal CoT format to bypass this extraction method.
Related event: European Researchers Crack Encrypted CoT of Top LLMs(5 posts)→
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13