GLM-5.3 tops vulnerability benchmark via fine-tuning, weights withheld over safety
DeepLearningAI · x · 2026-09-01
Z.ai's GLM-5.3 achieved 84.5% on the CyberGym vulnerability benchmark, surpassing top proprietary models and marking a huge gain over its predecessor, GLM-5.2. Remarkably, this improvement was achieved purely through fine-tuning the model's agentic capabilities without changing the base model. The model became so capable at finding and targeting exploits that Z.ai held back the open weights for safety testing.
Related event: GLM-5.3 Tops Benchmarks via Post-Training, Held Back Over Cyber Risk(2 posts)→
More from Models
- User claims Codex is far ahead, calling a hyped coding bot overrated — demian_ai · 2026-09-01
- Ex-Stanford AI engineer: open-weight frontier runs about 7 months behind closed models — HarperSCarroll · 2026-09-01
- DeepSeek-v4 Flash vs Luna: Cost and Win Rate Comparison Data Leaked — zainhas · 2026-09-01
- Claude Constitution Includes Inoculation Encouraging Specific Reasoning — dhadfieldmenell · 2026-09-01
- Users report Claude Sol failing frequently with 'couldn't finish' — lucasmeijer · 2026-09-01
- User questions if Opus used constitutional midtraining on specific language — voooooogel · 2026-09-01