New attack steals reasoning traces from OpenAI and Anthropic APIs

maksym_andr · x · 2026-08-16

A paper by Maksym Andriushchenko et al., "Stealing Reasoning Traces from Proprietary LLM APIs," reveals a critical vulnerability where encrypted reasoning blocks are compatible across different sessions and models. By injecting encrypted traces from a strong model into a weaker one, researchers forced it to decode the plaintext, successfully extracting private reasoning from Anthropic, OpenAI, and Google, and recovering 367 PII entries from public logs.

Original post →

More from Safety

Safety channel →