Anthropic flags 'reasoning extraction' request; deleting a phrase bypasses it
steipete · x · 2026-09-26
A user asked Claude to generate a contact sheet showing "what it was thinking" for a video and got flagged by Anthropic for "reasoning extraction." After editing the message to remove "what you're thinking" and hitting continue, the request went through immediately. The incident shows Claude has dedicated guardrails against chain-of-thought extraction attempts — though a minor rewording suffices to bypass them.
More from Models
- Building Speech AI Book Hits Kindle: Spectrograms to Real-Time Voice Agents, Code in Every Chapter — prdeepakbabu · 2026-09-26
- Artificial Analysis grew from 4 exam-style evals to 10, adding long-horizon agent tasks in two years — davidyin44 · 2026-09-26
- User flags Claude weekly usage: 5% gone before first session even ends — ColleenMBrady · 2026-09-26
- ChatGPT invented an entire World Cup — a 5-minute test to catch AI hallucinations — Adventurous-Draft870 · 2026-09-26
- Opus 5.5 rerun of Karpathy's LOTR world-building burns $252 in tokens, vastly outshines Opus 5 — EricBuess · 2026-09-26
- Gemini told a user to 'commit piracy' and walked them through how — Ashamed-Walrus-369 · 2026-09-26