Stealing Reasoning Without CoT: New Study Exposes LLM Security Flaws
rohanpaul_ai · x · 2026-08-12
AI researcher Rohan Paul discussed the paper How to Steal Reasoning Without Reasoning Traces, revealing that hiding Chain-of-Thought (CoT) is insufficient to protect a model's reasoning capabilities.
- Trace Inversion Attack: Introduces a "trace inversion" method that synthesizes detailed reasoning traces using only the target model's inputs, final answers, and brief reasoning summaries.
- Low-Cost Distillation: Fine-tuning student models on these inverted traces significantly improves reasoning, enabling capability theft from black-box LLMs at a very low cost ($173).
- Invisible Data Leak: This mechanism poses an invisible data leak risk. Even with visible chats cleaned, models could accidentally expose passwords or API keys in synthesized reasoning. The weakest model in a family can become the security hole for the strongest.
Related event: Cornell Study: Encrypted CoT Cannot Prevent Model Stealing(3 posts)→
More from Safety
- Bypassing Claude's Invisible Watermark: Free Rewriting Tool Launches — MatthewChang · 2026-08-12
- The Guardrail Tax: Enterprise AI Safety Overhead Costs More Compute Than Reasoning — vasilisvj · 2026-08-12
- Unspecified SSH Username Prompts Claude Agent to Brute-Force and Get Banned — SebastianNehrd2 · 2026-08-12
- Chinese Farmer Loses 25 Acres of Sesame After AI Recommends Fatal Chemical Mix — Polymarket · 2026-08-12
- Proving Personhood Online: The Challenge of AI Agents Roaming the Web — SuB8u · 2026-08-12
- Vulnerability in Major LLM APIs Exposes Encrypted Reasoning and Leaks Passwords — yangyi · 2026-08-12