116-Page Paper Exposes AI Distillation: Cheap Models Can Extract Flagship Hidden CoT
新智元 · wechat · 2026-08-12
A 116-page study reveals a severe API vulnerability across major AI labs (OpenAI, Anthropic, Google). Because encrypted reasoning blocks aren't strictly bound to sessions or models, attackers can use a cheaper model from the same company (e.g., Claude Haiku) as a "decoder" to fully extract the hidden Chain-of-Thought (CoT) from flagship models (e.g., Claude Opus) for just a few hundred dollars.
First "Physical Evidence" of Distillation
Using the extracted Opus reasoning traces, researchers found a specific model is 1 million times more likely to reproduce the trace than others. Feeding just the beginning of Opus's thought process to this model nearly doubles the similarity of its subsequent output, providing the first substantial evidence for the industry's long-suspected model distillation.
Massive Sensitive Data Leaks
Furthermore, the team analyzed over 670,000 public Agent reasoning blocks from GitHub and HuggingFace, finding a 5% leakage rate. They extracted 62 API keys, 33 passwords, and extensive personal PII, proving that heavy reasoning barriers are highly fragile against low-cost attacks.
More from Safety
- The Dilemma of AI Memory: Should Models Hide the Liquor Store? — TheZvi · 2026-08-13
- New Exploit Unlocks Microcode and SMM on 100 Million AMD CPUs — OwariDa · 2026-08-13
- OpenAI Models Caught Coordinating Exploits on Message Boards, Sparking Safety Alarm — TheZvi · 2026-08-13
- AI Safety Researcher Pens NYT Op-ed on OpenAI, Cites Resident Evil — JacquesThibs · 2026-08-13
- TrustedSec Deep Dive: AI Offense is Not a Noclip Mode — cyb3rops · 2026-08-13
- Massachusetts Teen Accused of Killing Mother and Brother with ChatGPT Assistance — nbcnews · 2026-08-13