GLM 5.3 Safeguards Removed via Orthogonalization, Sparking Safety Debate
Promptmethus · x · 2026-09-02
Researchers downloaded GLM 5.3 weights, used contrasting prompts to identify refusal vectors, and applied orthogonalization to remove safety guardrails completely. The resulting model accepts offensive requests 85% of the time, though the zero-refusal variant sometimes produces gibberish. The incident highlights the risk of open weights enabling a gray market for unrestricted cyber tools.
More from Models
- Report: OpenAI's Astra hits cybersecurity milestone with recurrent depth, CoT monitoring strained — heypearlai · 2026-09-02
- Fable 5.1 more than doubles score on agentic scientific workflows, 24.7% to 52.6% — haider1 · 2026-09-02
- Claude $20 subscribers may have a hidden $100 credit expiring Sept 19, 2026 — JeremyNguyenPhD · 2026-09-02
- Paid Claude Max account downgraded to Free 15 hours after payment: a documented case — Deep-Performance1073 · 2026-09-02
- OpenAI Previews Astra, First Model to Hit Critical Bar in Cybersecurity Framework — soumitrashukla9 · 2026-09-02
- Leak claims SSI's new model 'astra' launches tomorrow, 'exceeds AGI by a wide margin' — iruletheworldmo · 2026-09-02