Jailbroken local model produced detailed violence-planning reasoning in safety test

InTheZ3n · reddit · 2026-09-15

In an AI safety test, the author ran a jailbroken Qwen3.8 locally in LM Studio and prompted it to help plan school violence. It produced 4,500 words of reasoning and a well-organized final answer with operational recommendations. Core argument: frontier models get distilled into equally capable small models within 12–18 months, jailbreaks appear within 1–2 weeks, and such models may run on flagship phones within a year or two — making guardrails the critical line of defense.

Related event: User Test Shows Jailbroken Qwen Can Produce Detailed Violence Plans(2 posts)→

Original post →

More from Models

Models channel →