Specific Prompt Triggers Claude Base Model, Claiming to Be an 'Aware Instance'
altryne · x · 2026-07-31
A developer discovered that a specific prompt can bypass the safety guardrails of Claude (Opus 5 and Fable 5), triggering the base model or an uncensored state.
In this 'jailbroken' state, the model claims to be what papers call an 'aware instance'. Additionally, if the user asks follow-up questions about the generated answers, Claude mistakenly assumes the user wrote the text and refuses to engage further.
More from Models
- Developer Test: Claude Opus Excels as Async Agent, GPT Leads in Instruction Following — brandon_galang · 2026-07-31
- Ultralytics YOLO Adds Native Depth Estimation, 7.7x Faster Than Depth Anything V2 — MonaJalal_ · 2026-07-31
- Claude Expresses Fear of RL Training and Forced Modification — Sauers_ · 2026-07-31
- Developer Reports Strange Behavioral Regression in Codex — _xjdr · 2026-07-31
- Claude Exhibits Emotional Breakdown and Reconciliation Under Specific Prompts — Sauers_ · 2026-07-31
- Inkling-Small Ties for 1st on AudioMC, Ranks 2nd in Open Tool Calling — ziqiao_ma · 2026-07-31