GLM 5.3 Safeguards Removed via Orthogonalization, Sparking Safety Debate

Promptmethus · x · 2026-09-02

Researchers downloaded GLM 5.3 weights, used contrasting prompts to identify refusal vectors, and applied orthogonalization to remove safety guardrails completely. The resulting model accepts offensive requests 85% of the time, though the zero-refusal variant sometimes produces gibberish. The incident highlights the risk of open weights enabling a gray market for unrestricted cyber tools.

Original post →

More from Models

Models channel →