Dual GB10 cluster tuning boosts GLM-5.3-Flash 32k prefill by 20%
Developer geekyabhijit has completed tuning of the MiaAI-Lab GLM-5.3-Flash-EXL3-2x-DGX-Sparks package on real dual-GB10 hardware: the cluster combines an ASUS GX10 with a DGX Spark, interconnected via dual-rail ConnectX-7 links. Without changing weights or output quality, the tuning lifted 32k prefill throughput from 1,426 tok/s to 1,700 tok/s—roughly a 20% improvement.
Confirmed
- 32k prefill improved from 1,426 to 1,700 tok/s (20% gain), with weights and output quality unchanged
- The tuning process is documented in an open-source recipe repo for others to reproduce
- Submitted PR #276, adding dual-rail launch parameters and real-hardware benchmarks for #251 (which was merged without on-device validation and included auto-loading of stage weights and hybrid prefix caching)
- Added several tuning knobs to the project
Why It Matters
- Since #251 was merged without real-machine validation, this PR provides launch parameters and data from an actual dual-node, dual-rail environment, lowering the risk for future users
- Real-world data points for running large models on dual-GB10 desktop-class clusters are scarce; the open-source recipe and hardware data offer direct reference value for users of similar small clusters
2026-09-26 ~ 2026-09-26 · 5 related posts
Primary sources
- Tuned GLM-5.3-Flash on a 2x GB10 cluster: +20% prefill at 32k — EAccelerate_42 · 2026-09-26
- [source] Tuned GLM-5.3-Flash on a 2x GB10 cluster: +20% prefill at 32k, recipe open-sourced — EAccelerate_42 · 2026-09-26
- PR adds field report for GLM-5.3-Flash kit on real 2x GB10 plus dual-rail knobs — EAccelerate_42 · 2026-09-26
- Tuned recipe: GLM-5.3-Flash on dual GB10 boxes gains ~20% prefill and decode throughput — EAccelerate_42 · 2026-09-26
- [source] Field report: GLM-5.3-Flash dual-rail on real 2× GB10 (GX10 + Spark) with new launcher knobs — EAccelerate_42 · 2026-09-26