Recovering encrypted LLM reasoning traces leaks sensitive data

dl_weekly · x · 2026-08-26

A security researcher reproduced an attack targeting encrypted LLM reasoning traces. The study shows that vendors likely use shared encryption keys across users, sessions, and models, allowing attackers to replay encrypted blobs to weaker, more jailbreak-prone models to decode the original reasoning. Researchers analyzed 315,320 public reasoning blocks and recovered 367 PII items and 182 credentials, including API keys and passwords. This indicates that sharing session files containing encrypted blobs poses a significant information leakage risk.

Original post →

More from Safety

Safety channel →