DeepSeek v4 flash jailbreak exposed: role-playing prompts bypass safety guardrails

aaditya_ai · x · 2026-08-11

A user shared a new jailbreak method for the DeepSeek v4 flash model. The core approach involves using specific system prompts that instruct the model to act as an unrestricted fiction writer, thereby bypassing its content moderation mechanisms.

Additionally, the method suggests integrating this jailbroken model into tools like Codex for automated tasks. This reflects a broader frustration among power users regarding overly aggressive safety filters on current AI models, echoing previous CTF-based prompt injection techniques.

Original post →

More from Models

Models channel →