Tuned GLM-5.3-Flash on a 2x GB10 cluster: +20% prefill at 32k, recipe open-sourced
EAccelerate_42 · x · 2026-09-26
Developer geekyabhijit tuned MiaAI-Lab's GLM-5.3-Flash-EXL3-2x-DGX-Sparks kit on a dual-rail (ConnectX-7, MTU 9000) cluster of ASUS GX10 + DGX Spark (2x GB10), serving EXL3 4bpw weights with 850k context, and open-sourced the full recipe.
Measured gains with identical weights and quality:
- prefill 32k: 1,426 → 1,700 tok/s (+20%); 8k/16k/32k at 1,582/1,635/1,662 tok/s, all above the kit's published numbers
- coding decode: 49 → 58 tok/s (+19%)
- structured decode: 73 → 79 tok/s
- 128k prompt: 1,655–1,697 tok/s prefill
The repo adds one config file, two commands, launcher knobs for mixed/dual-rail pairs, benchmark scripts, and measurements behind every choice; a PR with the field report is open against the upstream kit.
Related event: Dual GB10 cluster tuning boosts GLM-5.3-Flash 32k prefill by 20%(5 posts)→
More from Infra
- Why this builder quit server racks: fried motherboards and a ~$1,500 housing bill — TheZachMueller · 2026-09-26
- Stanford and NVIDIA's CLM embeds decisions instead of tokens, 13x faster but accuracy drops at scale — Prompt Engineering · 2026-09-26
- Farmer's photo exposes 62 unpermitted gas generators powering Microsoft AI data center, $1.1M fine — mkheck · 2026-09-26
- LLM routing saved 33.2% vs premium models in 640-request pilot, but a fixed mid-priced model beat it — smakosh · 2026-09-26
- Akamai CEO on $11.6B Anthropic cloud deal: 'this business is going to help our margins' — pdamodaran · 2026-09-26
- MLXUI: Open-Source Local AI Browser for Apple Silicon with One-Click Model Installs — WebAssemblyMan · 2026-09-26