Attacks Restore Reasoning Traces, Exposing Proprietary LLM Chain-of-Thought

Machine Learning Street Talk · youtube · 2026-08-23

This episode features Ilia Shumailov and Alexander Panfilov discussing their paper on stealing reasoning traces from proprietary LLM APIs. The attack exploits a design flaw where providers return encrypted reasoning states for session resumption. Attackers can replay these blobs across users or models, using a smaller model to trick the provider into decrypting and revealing hidden reasoning in plain text. The discussion covers privacy leaks, reusable jailbreaks, poisoned agent traces, and potential defenses, advocating for controlled experiments over sweeping claims.

Original post →

More from Safety

Safety channel →