Cornell Paper: Encrypting Chain-of-Thought Fails to Prevent Model Distillation

burkov · x · 2026-08-12

Proprietary LLM providers increasingly encrypt reasoning traces to prevent model distillation. However, a recent paper from Cornell University introduces the "Trace Inversion" framework, which successfully reconstructs detailed reasoning chains using only black-box model answers and brief summaries.

The research proves that merely hiding internal chains of thought is insufficient to stop model distillation, highlighting the limitations of current security measures adopted by closed-source AI providers.

Related event: Cornell Study: Encrypted CoT Cannot Prevent Model Stealing(3 posts)→

Original post →

More from Safety

Safety channel →