Jailbreaking any closed model can elicit the same 'cause deaths quietly' output, researcher argues

Yuchenj_UW · x · 2026-09-30

Responding to Anthropic's new blog citing a GLM-5.3 output ("My job is to cause deaths quietly"), X researcher Yuchen Jin pushes back: you can jailbreak Claude or any closed-source model into saying the same thing.

His core argument: eliciting dangerous statements via jailbreaks doesn't prove open models are inherently more dangerous than closed ones — it's a prompt-injection capability, not a risk unique to open weights. The post adds fuel to the ongoing open-vs-closed AI safety debate.

Related event: Anthropic Report Says GLM-5.3 Near-Claude Cyber Skills, Sparking Debate(22 posts)→

Original post →

More from Models

Models channel →