Tuned GLM-5.3-Flash on a 2x GB10 cluster: +20% prefill at 32k

EAccelerate_42 · x · 2026-09-26

The author tuned MiaAI-Lab's GLM-5.3-Flash on an ASUS GX10 + DGX Spark dual-GB10 dual-rail cluster with identical weights and quality: 32k prefill rose from 1,426 to 1,700 tok/s (+20%), coding decode from 49 to 58 tok/s (+19%), structured decode from 73 to 79 tok/s, and 128k prompts hit 1,697 tok/s. Full config and measurements are in the open-source repo.

Related event: Dual GB10 cluster tuning boosts GLM-5.3-Flash 32k prefill by 20%(5 posts)→

Original post →

More from Infra

Infra channel →