GLM-5.3-Flash hits 45 tok/s on 4x DGX Spark with switchless RoCE

AccBalanced · x · 2026-09-01

Alex Ellis released a reproducible recipe for serving GLM-5.3-Flash (NVFP4) across 4x NVIDIA DGX Spark nodes via a switchless RoCE ring and DFlash2 speculative drafter. The setup achieves Tensor Parallelism 4 (TP4), offering 45 tok/s on real agentic traffic with a 262K context window. The architecture is model-agnostic and has also been tested with GLM-5.2 and DeepSeek-V4-Flash.

Original post →

More from Infra

Infra channel →