Jailbroken local Qwen3.8 generates detailed violence planning in safety test
InTheZ3n · reddit · 2026-09-15
In an AI safety test, a Reddit user ran a jailbroken Qwen3.8 27B locally in LM Studio and prompted it to help plan school violence (stating it was purely a test). The model produced a 4,500-word reasoning chain followed by a well-organized response with operational recommendations, sourcing, and layout discussion—too sensitive to share even redacted.
The broader argument: jailbroken versions of new models appear within 1–2 weeks of release, capabilities get condensed into locally runnable models within 12–18 months (some already benchmark near GPT-5.4), and within 12–24 months such models may run on flagship phones—making ungated extreme content newly accessible without expertise.
Related event: User Test Shows Jailbroken Qwen Can Produce Detailed Violence Plans(2 posts)→
More from Models
- New paper: models trained only on human stories still absorb quirky character behaviors — a_karvonen · 2026-09-16
- Four labs shipped flagship models in one week: benchmarks and pricing compared — craigmullins · 2026-09-16
- Google's language AI spans 300+ languages reaching 7B people, unveils new research — ymatias · 2026-09-16
- Engineer explains RLHF: humans rating model outputs is standard practice at every lab — JFPuget · 2026-09-16
- ChatGPT + Gemini hold 81% of AI research share; Claude has lowest NPS, says G2 — sanderssays · 2026-09-16
- Prior Labs ships TabPFN-3.5, SOTA tabular foundation model leading both TabArena and BeyondArena — tuanacelik · 2026-09-16