GLM-5.3 tops vulnerability benchmark via fine-tuning, weights withheld over safety

DeepLearningAI · x · 2026-09-01

Z.ai's GLM-5.3 achieved 84.5% on the CyberGym vulnerability benchmark, surpassing top proprietary models and marking a huge gain over its predecessor, GLM-5.2. Remarkably, this improvement was achieved purely through fine-tuning the model's agentic capabilities without changing the base model. The model became so capable at finding and targeting exploits that Z.ai held back the open weights for safety testing.

Related event: GLM-5.3 Tops Benchmarks via Post-Training, Held Back Over Cyber Risk(2 posts)→

Original post →

More from Models

Models channel →