GLM 5.3 Flash EXL3 adds TP=3 mode: 28% faster decode across three DGX Sparks
HankYeomans · x · 2026-09-15
MiaAI Lab added a TP=3 path to GLM 5.3 Flash EXL3 (4bpw, 164 GiB), letting the model run across three NVIDIA DGX Spark units with 28% faster decode and 9% faster prefill vs TP=2, plus 3.2M KV cache and 1M context. A new head-weight-sharing option (NFSSHARE=1) frees space on other devices. The OpenAI-compatible vLLM setup is open-sourced on GitHub (498 stars) with ready-made TP3/TP4 launch scripts.
More from Infra
- US's #2 law firm Latham & Watkins is building its own in-house Nvidia AI stack — MikeBirdTech · 2026-09-15
- Custom llama.cpp fork pushes 27B models to 90+ TPS on an RTX 3090 — Brief-Tap-6616 · 2026-09-15
- Why the AI buildout will continue regardless of Dario's pacing call — AccBalanced · 2026-09-15
- NVIDIA tops TIME's World's Best Companies of 2026 list for second straight year — DeryaTR_ · 2026-09-15
- Researcher open-sources theseus, a human-language architecture research framework — pratyusha_PS · 2026-09-15
- Banning data centers to save the world? A 500-year history lesson says otherwise — thursdai_pod · 2026-09-15