GLM 5.2 Flash Q2 Runs on M5 Max, Q4 Supports Tensor Parallel

antirez · x · 2026-08-28

antirez released GLM 5.2 Flash Q2, built with a recipe similar to DwarfStar GGUF, enabling it to run on an M5 Max with 128GB RAM. The Q4 weights also support tensor parallel inference across two M5 Max systems. The project is currently being tested and will be available on GitHub soon.

Related event: antirez Ports GLM 5.2 Flash to Run Locally on M5 Max(2 posts)→

Original post →

More from Infra

Infra channel →