GLM-5.2 Quantized Model Hits 16 tok/s on M3 Ultra 512GB

antirez · x · 2026-07-06

A user ran the GLM-5.2 q4 quantized model (ds4-eval version) on an Apple M3 Ultra 512GB device, achieving an inference speed of roughly 16 token/s. The developer plans to update older branches to introduce batch inference for further performance gains. This serves as a real-world benchmark for deploying Zhipu's GLM series models locally on high-end Apple silicon.

Related event: Quantized GLM-5.2 Runs Locally on M3 Ultra at 16 Tokens/s(2 posts)→

Original post →

More from Infra

Infra channel →