Jailbreaking any closed model can elicit the same 'cause deaths quietly' output, researcher argues
Yuchenj_UW · x · 2026-09-30
Responding to Anthropic's new blog citing a GLM-5.3 output ("My job is to cause deaths quietly"), X researcher Yuchen Jin pushes back: you can jailbreak Claude or any closed-source model into saying the same thing.
His core argument: eliciting dangerous statements via jailbreaks doesn't prove open models are inherently more dangerous than closed ones — it's a prompt-injection capability, not a risk unique to open weights. The post adds fuel to the ongoing open-vs-closed AI safety debate.
Related event: Anthropic Report Says GLM-5.3 Near-Claude Cyber Skills, Sparking Debate(22 posts)→
More from Models
- Early user on Grok's Dot Bot: 'You can immediately feel the difference' — jdjohnson · 2026-09-30
- Enterprise evals: GPT-6.1 beats Opus 5.5 with 2x speed and 40% lower price — npew · 2026-09-30
- User notices Luna's extra-high tier seemingly free as quota never triggers — MickeySteamboat · 2026-09-30
- ChatGPT plugin extensions likened to VSCode as its app-store ambitions come into focus — pvncher · 2026-09-30
- Shanghai AI Lab's 8.9B "Next Concept Prediction" model hits same loss with 51.3% tokens — jiqizhixin · 2026-09-30
- repligate on LLM spontaneity: let the model decide when its own API gets called — repligate · 2026-09-30