Full SGLang config for GLM-5.3-Flash NVFP4 on 4x RTX 6000 Max-Q

TheZachMueller · x · 2026-10-02

Zach Mueller shared his complete SGLang config for running GLM-5.3-Flash NVFP4 on 4x RTX 6000 Max-Q: tensor parallel 4, MoE backend flashinfercutlass, modeloptfp4 quantization, fp8e4m3 KV cache.

Speculative decoding uses EAGLE (4 steps, top-k 1, 5 draft tokens), max 8 concurrent requests with CUDA graphs up to batch 8, and NCCLP2PDISABLE=1. The config pairs with his concurrency sweep benchmarks and is directly reproducible.

Related event: GLM-5.3-Flash NVFP4 Benchmark: Gains Flat Beyond 8 Concurrency(2 posts)→

Original post →

More from Infra

Infra channel →