GLM-5.3 Series Shows Unusual Resistance to Abliteration Attacks
hunarbatra · x · 2026-08-29
Tests indicate the GLM-5.3 series is unusually resistant to "abliteration" attacks. Attempts to drive refusals to zero via more data, multi-directional subspaces, and layer-constrained ablation all failed. The hypothesis is that its safety policy is deeply distributed across weights, reinforced by SFT and preference/RL training, signaling impressive alignment work by the @Zaiorg team. Weights and a paper are coming soon.
More from Models
- Tencent compresses Hy4-preview from 1.5TB to 200GB GGUF — RedditUsr2 · 2026-08-29
- Zai releases GLM-5.3 as open-weight for agentic coding and cyber defense — alexcovo_eth · 2026-08-29
- Ornith 1.5 35B GGUF Model Received a Silent Update — miki4242 · 2026-08-29
- Community seeks best Minimax H3 Turbo LoRA settings and workflows — Enough-Bag-3891 · 2026-08-29
- MiniMax H3 Max Criticized for Extreme Censorship on Fal — MrUtterNonsense · 2026-08-29
- User Predicts Qwen-4-27B Will Be a Game Changer — Steus_au · 2026-08-29