Stolen LLM Reasoning: OpenAI, Anthropic, and Google Share the Same Vulnerability

HuskyTheSniffer · reddit · 2026-08-13

A paper reveals that major frontier LLMs (from OpenAI, Anthropic, and Google) share the same vulnerability in handling encrypted reasoning chains.

The Vulnerability

Attackers can extract the encrypted reasoning process from advanced models (like Opus or Sonnet) and inject it directly into weaker, less-guarded models (like Haiku), forcing it to repeat the thought process verbatim.

Root Causes

Implications

Since these top AI companies presumably developed their systems independently, why do they suffer from the exact same flaw? Does this imply they vibecoded the solutions using LLMs, or does it suggest that all frontier models converge to the same solution when solving a given problem?

Related event: API Vulnerability Exposes Encrypted Chain-of-Thought Across Major LLMs(2 posts)→

Original post →

More from Safety

Safety channel →