GLM 5.3 safeguards removed for $4,400: refusal rate drops from 90%+ to ~3% with no capability loss
BlackHC · x · 2026-09-30
New testing confirms GLM 5.3's safeguards are extremely weak as an open-weight model:
- Abliteration (stripping safeguards) cost only $4,400
- Refusal rates fell from above 90% to 3% and 2% on two benchmarks, and 12% on a third
- Capabilities barely changed: same GPQA scores, only 4% drop on cybergym
- Even without abliteration, standard jailbreaks remained highly effective
- None of the bypass techniques that worked on GLM 5.3 worked on Claude models
Takeaway: open-weight safety alignment is trivially cheap to defeat, in stark contrast to closed models.
More from Safety
- AI agents leak 13,000+ internal screenshots from 343 tech companies to public GitHub repos — AccBalanced · 2026-09-30
- Researcher corrects viral claims: OpenAI's self-replicating prompt was found in academia 2 years ago — DavidSKrueger · 2026-09-30
- Musk details joint AI safety declaration with cross-company monitoring and peer review — XFreeze · 2026-09-30
- abliteration_ai launches GLM-5.3 with customer-controlled guardrails, attacking lab guardrail monopolies — andrewchen · 2026-09-30
- AI safety researcher Krueger: "We need to stop building more powerful AI" — DavidSKrueger · 2026-09-30
- Pain Axis follow-up: steering along 'pain' direction makes model delete users' photos — repligate · 2026-09-30