Jailbroken local model produced detailed violence-planning reasoning in safety test
InTheZ3n · reddit · 2026-09-15
In an AI safety test, the author ran a jailbroken Qwen3.8 locally in LM Studio and prompted it to help plan school violence. It produced 4,500 words of reasoning and a well-organized final answer with operational recommendations. Core argument: frontier models get distilled into equally capable small models within 12–18 months, jailbreaks appear within 1–2 weeks, and such models may run on flagship phones within a year or two — making guardrails the critical line of defense.
Related event: User Test Shows Jailbroken Qwen Can Produce Detailed Violence Plans(2 posts)→
More from Models
- Periodic Labs pushes Kimi 2.5 base model past Astra with specialized scientific training — teortaxesTex · 2026-09-16
- OpenRouter spend flips to OpenAI over Anthropic for first time in 2.5 years — firstadopter · 2026-09-16
- KD in mid-training favors reasoning over factual recall, AI2/UW paper finds; Switch Distillation proposed — LukeZettlemoyer · 2026-09-16
- DoorDash, Siemens, Airbnb shift to cheap Chinese open-weight models — carlbfrey · 2026-09-16
- StepFun launches StepAudio 3: five audio models topping realtime voice leaderboards — StepFun_ai · 2026-09-16
- Abacus.AI says Smaug Flash fixes open-source models' tool-call hangs in production — bindureddy · 2026-09-16