GLM 5.2 Flash Q2 Runs on M5 Max, Q4 Supports Tensor Parallel
antirez · x · 2026-08-28
antirez released GLM 5.2 Flash Q2, built with a recipe similar to DwarfStar GGUF, enabling it to run on an M5 Max with 128GB RAM. The Q4 weights also support tensor parallel inference across two M5 Max systems. The project is currently being tested and will be available on GitHub soon.
Related event: antirez Ports GLM 5.2 Flash to Run Locally on M5 Max(2 posts)→
More from Infra
- AI Energy Crisis: Data Centers to Exceed Germany's Usage by 2030 — ingliguori · 2026-08-28
- AMD Releases ROCm 10: AI-Native Stack with 3.3x Inference Boost — AnushElangovan · 2026-08-28
- Ricursive Accelerates Chip Design Stage 1000x Using AI — annadgoldie · 2026-08-28
- Perplexity Launches Portable Computer: Fully Local, Cloud-Free AI Agent — AravSrinivas · 2026-08-28
- Proposal: Hardcoding AI Models into SoCs for Instant Inference — kingslayerer · 2026-08-28
- Locus Pro Launches Self-Serve as 'OpenRouter for Everything' with 3,300+ Providers — kleffew94 · 2026-08-28