Tuned recipe: GLM-5.3-Flash on dual GB10 boxes gains ~20% prefill and decode throughput

EAccelerate_42 · x · 2026-09-26

EAccelerate42 tuned MiaAI-Lab's GLM-5.3-Flash (EXL3 4bpw, 850k context) serving kit on a 2× GB10 cluster (ASUS GX10 + DGX Spark, dual-rail) and open-sourced the measured recipe—same weights, same output quality:

Config details: TP=2, both ConnectX-7 rails at MTU 9000, driver 580.173.02, benched with unique-salt cold prompts, thinking off, temp 0, median of 5×400 tokens. The repo (e-accelerate/glm53-flash-2x-recipe) ships patches, launcher scripts, adaptivek config, and measurements behind every setting—a directly reproducible dual-box deployment recipe.

Related event: Dual GB10 cluster tuning boosts GLM-5.3-Flash 32k prefill by 20%(5 posts)→

Original post →

More from Infra

Infra channel →