Tuned recipe: GLM-5.3-Flash on dual GB10 boxes gains ~20% prefill and decode throughput
EAccelerate_42 · x · 2026-09-26
EAccelerate42 tuned MiaAI-Lab's GLM-5.3-Flash (EXL3 4bpw, 850k context) serving kit on a 2× GB10 cluster (ASUS GX10 + DGX Spark, dual-rail) and open-sourced the measured recipe—same weights, same output quality:
- prefill @32k: 1,426 → 1,700 tok/s (+20%)
- coding decode: 49 → 58 tok/s (+19%)
- structured decode: 73 → 79 tok/s
- 128k prompt: 1,697 tok/s
Config details: TP=2, both ConnectX-7 rails at MTU 9000, driver 580.173.02, benched with unique-salt cold prompts, thinking off, temp 0, median of 5×400 tokens. The repo (e-accelerate/glm53-flash-2x-recipe) ships patches, launcher scripts, adaptivek config, and measurements behind every setting—a directly reproducible dual-box deployment recipe.
Related event: Dual GB10 cluster tuning boosts GLM-5.3-Flash 32k prefill by 20%(5 posts)→
More from Infra
- A100 SXM4 connector production ends, forcing new tooling investment — EAccelerate_42 · 2026-09-26
- Two local AIs talk to each other with no cloud: hands-on with Braid — Scobleizer · 2026-09-26
- The handiest GPU this dev ever bought is a ~$300 Intel Arc A310, not NVIDIA or AMD — TheZachMueller · 2026-09-26
- Running image generation in the browser on local hardware: 10s pixel art on an RTX 3060 — Bartholomheow · 2026-09-26
- Alibaba's T-Head unveils Zhenwu V900 chip: 216GB per card, sales in Q1 2027 — shashib · 2026-09-26
- llama.cpp fork dedups repeated prompts losslessly, cutting 108k to 71k tokens in agent loops — Odd_Cauliflower_8004 · 2026-09-26