116-Page Paper Reveals Vulnerability: Extracting Encrypted Reasoning from Top LLMs

xeophon · x · 2026-08-11

A new 116-page paper highlights a significant vulnerability in models from OpenAI, Anthropic, and Gemini, allowing attackers to extract encrypted raw reasoning at scale.

This vulnerability leads to multiple security issues, including distillation attacks and credential extraction. The researchers also found numerous instances of illegible reasoning (especially in GPT models), unfaithful reasoning, and evidence suggesting certain open-weight models were likely distilled.

Related event: Study Reveals API Flaw to Extract Encrypted Reasoning Traces and Evidence of Distillation(38 posts)→

Original post →

More from Safety

Safety channel →