GLM 5.3 Flash EXL3 adds TP=3 mode: 28% faster decode across three DGX Sparks

HankYeomans · x · 2026-09-15

MiaAI Lab added a TP=3 path to GLM 5.3 Flash EXL3 (4bpw, 164 GiB), letting the model run across three NVIDIA DGX Spark units with 28% faster decode and 9% faster prefill vs TP=2, plus 3.2M KV cache and 1M context. A new head-weight-sharing option (NFSSHARE=1) frees space on other devices. The OpenAI-compatible vLLM setup is open-sourced on GitHub (498 stars) with ready-made TP3/TP4 launch scripts.

Original post →

More from Infra

Infra channel →