GLM-5.3 cyber eval: 75% CVE rediscovery beats GPT-5.6, costs 40% less
joshua_saxe · x · 2026-08-15
A third-party security evaluation of GLM-5.3 shows it rediscovered 75% of CVEs at pass@3, 7 points higher than GPT-5.6-Terra, while costing 40% less to run. At pass@1, it surfaces 60.4% of CVEs on average, the highest for an open-source model. It also reports fewer false positives than DeepSeek v4 models, making its output more trustworthy. The evaluator calls GLM-5.3 the strongest and most consistent OSS model for cyber tasks.
Related event: Zhipu Releases GLM-5.3 Model(37 posts)→
More from Models
- Anthropic uses internal model far better than Mythos 5, no release planned — kimmonismus · 2026-08-15
- Test shows Qwen3.8-27B achieves high precision in long video timestamping — vanstriendaniel · 2026-08-15
- RTX 3090 gets 35 t/s on Qwen 3.8 27B — cviperr33 · 2026-08-15
- Gemini 3.7 Flash Launches with Enhanced Reasoning and Tool Use — arvind_io · 2026-08-15
- Anthropic confirms Mythos 5 as top internal model, hints at mysterious Model 2 — scaling01 · 2026-08-15
- Users report random stop behavior in Qwen 2.5/3 during long context generation — T_rex2700 · 2026-08-15