DwarfStar Adds GLM 5.3 Flash Support with Q2/Q4 on MacBook
antirez · x · 2026-08-28
The DwarfStar project introduced a glm-5.3-flash branch, supporting Q2 and Q4 quantization for GLM 5.3 Flash. It enables inference on a single 128GB MacBook or tensor parallel inference across two 128GB MacBooks via RDMA, achieving 37 t/s generation and 500 t/s prefill. DGX Spark is supported, with ROCm coming soon.
More from Infra
- Optimizing Minimax H3 Inference Speeds on Consumer Hardware — Ambitious_Fold_2874 · 2026-08-28
- Acorn: AI Chief of Staff running on your own server — MatthewChang · 2026-08-28
- Trade unions push back against data center opposition — MatthewBerman · 2026-08-28
- UK Labour rejects Green party call to halt AI datacentre construction — nordicinst · 2026-08-28
- llama.cpp mmap fits Qwen3.8-Flash-Next in 16G+64G RAM at 26t/s — q8019222 · 2026-08-28
- x402: Internet-Native Payment Standard for AI Agents — kleffew94 · 2026-08-28