antirez ports GLM 5.2 Flash to run on M5 Max, TP across two Macs
antirez · x · 2026-08-28
antirez released GLM 5.2 Flash Q2 and Q4 quantizations, created with a recipe similar to DS4 Flash. It runs on M5 Max with 128GB RAM, and Q4 weights support tensor parallel inference across two M5 Max systems. Testing on CUDA and ROCm is ongoing; code coming to GitHub soon.
Related event: antirez Ports GLM 5.2 Flash to Run Locally on M5 Max(2 posts)→
More from Infra
- ThursdAI episode: Andy Masley actually checks the water and energy math on data centers — altryne · 2026-08-28
- AISI releases open-source optstop to reduce token costs in LLM evaluations — HZoete · 2026-08-28
- Nvidia projects ~70% revenue growth next fiscal year, driven by surging AI demand — Polymarket · 2026-08-28
- GPU Gold Rush: Are Tech Companies Creating a Hype-Driven Bubble? — DavidLinthicum · 2026-08-28
- Engineering Discussion: Optimizing Context Compaction in DeepSeek Harness — shady101852 · 2026-08-28
- NVIDIA Vera CPU ships at scale to accelerate agentic workloads — rohanpaul_ai · 2026-08-28