GLM-5.3-Flash hits 45 tok/s on 4x DGX Spark with switchless RoCE
AccBalanced · x · 2026-09-01
Alex Ellis released a reproducible recipe for serving GLM-5.3-Flash (NVFP4) across 4x NVIDIA DGX Spark nodes via a switchless RoCE ring and DFlash2 speculative drafter. The setup achieves Tensor Parallelism 4 (TP4), offering 45 tok/s on real agentic traffic with a 262K context window. The architecture is model-agnostic and has also been tested with GLM-5.2 and DeepSeek-V4-Flash.
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- mlx-signal-processing brings 10-200x faster signal ops to Apple Silicon — TheMoonMidas · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01