GLM 5.2 Quantized Runs Locally at 16 tokens/s on M3 Ultra
antirez · x · 2026-07-05
Ivan Fioravanti shared a video demo showing that a 4bit quantized version of Zhipu's GLM 5.2 can run locally at about 16 tokens/second on a single Apple M3 Ultra equipped with 512GB of RAM, while the ds4-eval video runs on the q2 quantized version. The demo was reshared by Redis creator antirez, showcasing the real-world performance of running large models locally on high-memory Macs.
Related event: Quantized GLM-5.2 Runs Locally on M3 Ultra at 16 Tokens/s(2 posts)→
More from Infra
- WSJ says Nvidia is weighing a $250B debt guarantee for OpenAI’s Ohio data center lease — mkheck · 2026-07-27
- Nvidia reportedly weighs $250B financing backstop for OpenAI’s Ohio data center — AccBalanced · 2026-07-27
- WEKA NeuralMesh is said to match HBM3 bandwidth on GPU servers — AccBalanced · 2026-07-27
- Nvidia gets mocked as “the leading open-source AI company” while repo chart shows it ahead — AccBalanced · 2026-07-27
- $8 ESP32-S3 runs a 28.9M-parameter LLM fully offline at 9.5 tokens per second — yangyi · 2026-07-27
- YC talk on BCI x AI says infrastructure is what really determines speed — garrytan · 2026-07-27