Anthropic Report Says GLM-5.3 Safeguards Bypassed 64%-100% of the Time, Drawing Open-Source Backlash

robleclerc · x · 2026-09-30

Anthropic 发布报告称,智谱最新模型 GLM-5.3 具备与 Claude Mythos Preview 相当的自主构建端到端网络攻击利用的能力,但缺乏有效安全护栏:在模拟测试中,攻击者用简单技术绕过 GLM-5.3 护栏的成功率为 64% 到 100%,而受护栏保护的 Claude 模型未被攻破。NIST 此前也评定 GLM-5.3 为「迄今网络能力最强的开源权重模型」。Rob LeCleerc 转发并批评 Anthropic 以「安全」之名攻击开源权重模型,称「恐惧是威权者的武器」,引发开源安全与模型护栏强度的争议。

Related event: Anthropic Report: GLM-5.3's Cyber Attack Capabilities Near Claude's, Guardrails Easily Bypassed(15 posts)→

Original post →

More from Models

Models channel →