116-Page Paper Reveals Vulnerability: Extracting Encrypted Reasoning from Top LLMs
xeophon · x · 2026-08-11
A new 116-page paper highlights a significant vulnerability in models from OpenAI, Anthropic, and Gemini, allowing attackers to extract encrypted raw reasoning at scale.
This vulnerability leads to multiple security issues, including distillation attacks and credential extraction. The researchers also found numerous instances of illegible reasoning (especially in GPT models), unfaithful reasoning, and evidence suggesting certain open-weight models were likely distilled.
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13