DwarfStar Adds GLM 5.3 Flash Support with Q2/Q4 on MacBook

antirez · x · 2026-08-28

The DwarfStar project introduced a glm-5.3-flash branch, supporting Q2 and Q4 quantization for GLM 5.3 Flash. It enables inference on a single 128GB MacBook or tensor parallel inference across two 128GB MacBooks via RDMA, achieving 37 t/s generation and 500 t/s prefill. DGX Spark is supported, with ROCm coming soon.

Original post →

More from Infra

Infra channel →