GLM 5.2 Quantized Runs Locally at 16 tokens/s on M3 Ultra

antirez · x · 2026-07-05

Ivan Fioravanti shared a video demo showing that a 4bit quantized version of Zhipu's GLM 5.2 can run locally at about 16 tokens/second on a single Apple M3 Ultra equipped with 512GB of RAM, while the ds4-eval video runs on the q2 quantized version. The demo was reshared by Redis creator antirez, showcasing the real-world performance of running large models locally on high-memory Macs.

Related event: Quantized GLM-5.2 Runs Locally on M3 Ultra at 16 Tokens/s(2 posts)→

Original post →

More from Infra

Infra channel →