Encrypted Chain-of-Thought in Proprietary LLMs Can Be Extracted via Weaker Sibling Models

Simon Willison · rss · 2026-08-12

A recent paper highlights a severe security vulnerability in the encrypted chain-of-thought (CoT) returned by major proprietary LLM APIs (OpenAI, Anthropic, Google).

The Vulnerability

Implications & Novel Prompt Injection

All affected vendors have acknowledged the report and subsequently patched the issue, rendering the attacks unsuccessful now.

Original post →

More from Models

Models channel →