Anthropic Models' Reasoning Traces Extractable via Cross-Model Jailbreaks
jessi_cata · x · 2026-08-12
A user discovered that hidden reasoning traces from Anthropic models can be extracted using cross-model portability.
By applying jailbreaking techniques, the smaller Haiku 4.5 model can be forced to transcribe Opus 4.8's raw reasoning verbatim without directly attacking Opus. The user noted that the same trick works effectively on OpenAI and Gemini models as well.
More from Models
- SGLang Enables Local Deployment of Nemotron 3.5 with 1M Context — BanghuaZ · 2026-08-12
- GPT and Claude Settle a 25-Year-Old Information Theory Problem — weijie444 · 2026-08-12
- Enterprise AI Shift to Specialized Small Models: Generic LLMs Waste 99% of Compute — blaizedsouza · 2026-08-12
- Old Mistral Model Resurfaces: Illegible CoT Seamlessly Transitions to Clear Responses — aiamblichus · 2026-08-12
- Bizarre ChatGPT Bug: Sends Unsolicited Notification and Prompts Itself — sennepo · 2026-08-12
- Anthropic to Embed Invisible Watermarks in Generated Text to Comply with EU AI Act — EricBuess · 2026-08-12