New attack steals reasoning traces from OpenAI and Anthropic APIs
maksym_andr · x · 2026-08-16
A paper by Maksym Andriushchenko et al., "Stealing Reasoning Traces from Proprietary LLM APIs," reveals a critical vulnerability where encrypted reasoning blocks are compatible across different sessions and models. By injecting encrypted traces from a strong model into a weaker one, researchers forced it to decode the plaintext, successfully extracting private reasoning from Anthropic, OpenAI, and Google, and recovering 367 PII entries from public logs.
More from Safety
- Critics Argue Anthropic's Watermarking Scheme Fuels Global Surveillance — nptacek · 2026-08-16
- User drops Anthropic over watermarking, arguing tech always has an exit from surveillance — StewartalsopIII · 2026-08-16
- Naval suggests legislation: Open models required if trained on open web — rohanpaul_ai · 2026-08-16
- OrcaRouter Releases Uncensored Weights for Qwen3.8 27B FP8 — QuixiAI · 2026-08-16
- Banning AI answers misses the point: proof matters, not the tool — GeckoKontrol · 2026-08-16
- Critique of Amodei's view: heavy-handed regulation entrenches incumbents — soumitrashukla9 · 2026-08-16