New Research: Hidden Reasoning Can Be Extracted from Encrypted Traces of Proprietary LLMs
daniel_mac8 · x · 2026-08-11
A new study reveals that encrypted chain-of-thought blocks returned by frontier models like Anthropic, OpenAI, and Google can be replayed and exploited to extract hidden reasoning. By replaying a strong model's trace into a weaker sibling and jailbreaking it, attackers recover the stronger model's reasoning in plaintext without directly attacking it or triggering anti-distillation safeguards. The research includes a paper and an interactive game, demonstrating extraction in just two API calls.
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13