GLM-5.3 Series Shows Unusual Resistance to Abliteration Attacks

hunarbatra · x · 2026-08-29

Tests indicate the GLM-5.3 series is unusually resistant to "abliteration" attacks. Attempts to drive refusals to zero via more data, multi-directional subspaces, and layer-constrained ablation all failed. The hypothesis is that its safety policy is deeply distributed across weights, reinforced by SFT and preference/RL training, signaling impressive alignment work by the @Zaiorg team. Weights and a paper are coming soon.

Original post →

More from Models

Models channel →