Attacks Restore Reasoning Traces, Exposing Proprietary LLM Chain-of-Thought
Machine Learning Street Talk · youtube · 2026-08-23
This episode features Ilia Shumailov and Alexander Panfilov discussing their paper on stealing reasoning traces from proprietary LLM APIs. The attack exploits a design flaw where providers return encrypted reasoning states for session resumption. Attackers can replay these blobs across users or models, using a smaller model to trick the provider into decrypting and revealing hidden reasoning in plain text. The discussion covers privacy leaks, reusable jailbreaks, poisoned agent traces, and potential defenses, advocating for controlled experiments over sweeping claims.
More from Safety
- Scholar uses AI-hallucinated citations, blames Google Scholar search results — aniketapanjwani · 2026-08-24
- AI Privacy and Safety Intervention: Who Defines the Boundaries of "Danger"? — Radiant_Dragonfruit3 · 2026-08-24
- Prediction market gives 68% chance of a state data center moratorium by year-end — Polymarket · 2026-08-24
- Blogger plans to explore link between data center backlash and AI safety — AndyMasley · 2026-08-24
- Richard Ngo to release part 2 of retrospective on alignment failures — RichardMCNgo · 2026-08-24
- US AI Data Centers Spark Death Threats; GOP Sees Electoral Risk — rohanpaul_ai · 2026-08-24