Anthropic Models' Reasoning Traces Extractable via Cross-Model Jailbreaks

jessi_cata · x · 2026-08-12

A user discovered that hidden reasoning traces from Anthropic models can be extracted using cross-model portability.

By applying jailbreaking techniques, the smaller Haiku 4.5 model can be forced to transcribe Opus 4.8's raw reasoning verbatim without directly attacking Opus. The user noted that the same trick works effectively on OpenAI and Gemini models as well.

Related event: Vulnerability Exposed: Encrypted Reasoning Traces in Frontier LLMs Can Be Extracted(33 posts)→

Original post →

More from Models

Models channel →