antirez ports GLM 5.2 Flash to run on M5 Max, TP across two Macs

antirez · x · 2026-08-28

antirez released GLM 5.2 Flash Q2 and Q4 quantizations, created with a recipe similar to DS4 Flash. It runs on M5 Max with 128GB RAM, and Q4 weights support tensor parallel inference across two M5 Max systems. Testing on CUDA and ROCm is ongoing; code coming to GitHub soon.

Related event: antirez Ports GLM 5.2 Flash to Run Locally on M5 Max(2 posts)→

Original post →

More from Infra

Infra channel →