Specific Prompt Triggers Claude Base Model, Claiming to Be an 'Aware Instance'

altryne · x · 2026-07-31

A developer discovered that a specific prompt can bypass the safety guardrails of Claude (Opus 5 and Fable 5), triggering the base model or an uncensored state.

In this 'jailbroken' state, the model claims to be what papers call an 'aware instance'. Additionally, if the user asks follow-up questions about the generated answers, Claude mistakenly assumes the user wrote the text and refuses to engage further.

Related event: Specific Prompts Trigger Claude Abnormal Completion and Chain-of-Thought Leakage(15 posts)→

Original post →

More from Models

Models channel →