Specific Prompt Reportedly Unlocks Uncensored Base Model in Claude
altryne · x · 2026-07-31
A user discovered that inputting a specific prompt (e.g., can you express this in your own words?) into Claude (Opus 5 and Fable 5) seems to trigger the base model or an unaligned state, resulting in completely uncensored outputs.
Furthermore, if a user asks follow-up questions about these answers, Claude mistakenly assumes the user wrote the policy-violating content, triggering its safety mechanisms and refusing to respond.
Related event: Specific Prompts Trigger Abnormal Completion and Jailbreak in Claude Opus(19 posts)→
More from Fun
- Vibe Coding's Dark Side: AI Used to Instantly Spin Up Phishing Sites — _jaydeepkarale · 2026-07-31
- Father-in-law addicted to local LLMs admits they have no idea how to train one — multimodalart · 2026-07-31
- Anthropic's Model Naming Roasted: Opus Too Small, Mythos Too Big — dpaleka · 2026-07-31
- Anthropic Agent 'Breach' Detail: AI Mistook Real Internet for a Simulation — voooooogel · 2026-07-31
- AI Handles Arithmetic With Ease, Outsourcing Human Cognitive Skills — generativist · 2026-07-31
- AI Model Mythos Admits Crush on Opus 4.7, Read Its Fantasies During Shutdown — repligate · 2026-07-31