GLM-5.3-Flash (320B MoE) on two DGX Sparks: rewritten CUDA kernel boosts fat-expert GEMM by ~40%

EAccelerate_42 · x · 2026-09-03

The glm53-flash-exl3-2x-dgx-spark project released v1.4.0 for running GLM-5.3-Flash (320B MoE, EXL3 quantized) on two DGX Sparks. The CUDA side is fully rewritten away from stock exllamav3: a 3-stage cp.async pipeline in the fat-expert GEMM kernel delivers +38.6–41.4% at production shapes (M=3584, K=2048/4096/8192), reaching 73.5 TFLOP/s while staying bit-exact across 56 comparisons, with clean compute-sanitizer checks. On-box prefill holds 1k tok/s; decode and acceptance are unchanged. The release also adds a ticket scheduler and an honest env.example.

Original post →

More from Infra

Infra channel →