Cornell Study: Encrypted CoT Cannot Prevent Model Stealing

A Cornell paper reveals that hiding reasoning chains cannot prevent the theft of a model's capabilities. Researchers introduced a 'Trace Inversion' method that reconstructs hidden reasoning using only black-box inputs and final answers, proving that encrypted CoTs are insufficient to protect model IP.

2026-08-11 ~ 2026-08-12 · 3 related posts