Encrypted Reasoning in Closed Models is 100% Recoverable, Paper Shows

Dany0 · reddit · 2026-08-11

A Reddit user shared a paper indicating that the encrypted reasoning processes of closed-source models like OpenAI's are 100% recoverable. The author urged the community to upload millions of traces from models like Anthropic's Claude to Hugging Face before providers patch the exploit.

Related event: Study Reveals API Flaw to Extract Encrypted Reasoning Traces and Evidence of Distillation(38 posts)→

Original post →

More from Safety

Safety channel →