Anthropic Report Says GLM-5.3 Safeguards Bypassed 64%-100% of the Time, Drawing Open-Source Backlash
robleclerc · x · 2026-09-30
Anthropic 发布报告称,智谱最新模型 GLM-5.3 具备与 Claude Mythos Preview 相当的自主构建端到端网络攻击利用的能力,但缺乏有效安全护栏:在模拟测试中,攻击者用简单技术绕过 GLM-5.3 护栏的成功率为 64% 到 100%,而受护栏保护的 Claude 模型未被攻破。NIST 此前也评定 GLM-5.3 为「迄今网络能力最强的开源权重模型」。Rob LeCleerc 转发并批评 Anthropic 以「安全」之名攻击开源权重模型,称「恐惧是威权者的武器」,引发开源安全与模型护栏强度的争议。
More from Models
- Claude Sonnet 5.5 coming to LMArena for limited-time testing — arena · 2026-09-30
- ARC Prize to Evaluate DeepSeek V4.1 Flash After Predecessor Hit 61.4% on ARC-AGI-2 — teortaxesTex · 2026-09-30
- gpt-6.1-sol grinds 35+ minutes on trivial validation prompt at xhigh setting — arthurcolle · 2026-09-30
- ChatGPT Pro users report 6-Pro web chats capped at 100 per week — triestdain · 2026-09-30
- GPT-6.1 Sol fixes 44 of 105 planted bugs for $6.56, matching Astra at a fraction of the cost — PawelHuryn · 2026-09-30
- 5 months after Mythos Preview panic, an open model already matches it — mariofilhoml · 2026-09-30