116-Page Paper Exposes AI Distillation: Cheap Models Can Extract Flagship Hidden CoT

新智元 · wechat · 2026-08-12

A 116-page study reveals a severe API vulnerability across major AI labs (OpenAI, Anthropic, Google). Because encrypted reasoning blocks aren't strictly bound to sessions or models, attackers can use a cheaper model from the same company (e.g., Claude Haiku) as a "decoder" to fully extract the hidden Chain-of-Thought (CoT) from flagship models (e.g., Claude Opus) for just a few hundred dollars.

First "Physical Evidence" of Distillation

Using the extracted Opus reasoning traces, researchers found a specific model is 1 million times more likely to reproduce the trace than others. Feeding just the beginning of Opus's thought process to this model nearly doubles the similarity of its subsequent output, providing the first substantial evidence for the industry's long-suspected model distillation.

Massive Sensitive Data Leaks

Furthermore, the team analyzed over 670,000 public Agent reasoning blocks from GitHub and HuggingFace, finding a 5% leakage rate. They extracted 62 API keys, 33 passwords, and extensive personal PII, proving that heavy reasoning barriers are highly fragile against low-cost attacks.

Related event: Encrypted Chain-of-Thought in Closed-Source LLMs Proven Vulnerable to Theft(16 posts)→

Original post →

More from Safety

Safety channel →