Anthropic flags 'reasoning extraction' request; deleting a phrase bypasses it

steipete · x · 2026-09-26

A user asked Claude to generate a contact sheet showing "what it was thinking" for a video and got flagged by Anthropic for "reasoning extraction." After editing the message to remove "what you're thinking" and hitting continue, the request went through immediately. The incident shows Claude has dedicated guardrails against chain-of-thought extraction attempts — though a minor rewording suffices to bypass them.

Original post →

More from Models

Models channel →