GLM-5.3-Flash (320B MoE) on two DGX Sparks: rewritten CUDA kernel boosts fat-expert GEMM by ~40%
EAccelerate_42 · x · 2026-09-03
The glm53-flash-exl3-2x-dgx-spark project released v1.4.0 for running GLM-5.3-Flash (320B MoE, EXL3 quantized) on two DGX Sparks. The CUDA side is fully rewritten away from stock exllamav3: a 3-stage cp.async pipeline in the fat-expert GEMM kernel delivers +38.6–41.4% at production shapes (M=3584, K=2048/4096/8192), reaching 73.5 TFLOP/s while staying bit-exact across 56 comparisons, with clean compute-sanitizer checks. On-box prefill holds 1k tok/s; decode and acceptance are unchanged. The release also adds a ticket scheduler and an honest env.example.
More from Infra
- fchollet posts a #TPU meme that leaves Hesamation exclaiming 'BRO WHAT?' — algo_diver · 2026-09-03
- Nvidia's CoWoS share seen falling to 45.6% by 2027 as AMD doubles to 19.8% — tengyanAI · 2026-09-03
- Loudoun County's 20-year data center history previews America's AI infrastructure future — suchenzang · 2026-09-03
- Alexandr Wang mocks NIMBYs: 'Planes come and go, but the data center hum never stops' — suchenzang · 2026-09-03
- Beff Jezos: Biology's Compute-per-Watt Is Massively Underestimated, Bio-Silicon Complexification Begins — beffjezos · 2026-09-03
- AMD's official Qwen3.8 MXFP4 quant fails to load in llama.cpp, is support still missing? — mailto_devnull · 2026-09-03