Tuned GLM-5.3-Flash on a 2x GB10 cluster: +20% prefill at 32k
EAccelerate_42 · x · 2026-09-26
The author tuned MiaAI-Lab's GLM-5.3-Flash on an ASUS GX10 + DGX Spark dual-GB10 dual-rail cluster with identical weights and quality: 32k prefill rose from 1,426 to 1,700 tok/s (+20%), coding decode from 49 to 58 tok/s (+19%), structured decode from 73 to 79 tok/s, and 128k prompts hit 1,697 tok/s. Full config and measurements are in the open-source repo.
Related event: Dual GB10 cluster tuning boosts GLM-5.3-Flash 32k prefill by 20%(5 posts)→
More from Infra
- A100 SXM4 connector production ends, forcing new tooling investment — EAccelerate_42 · 2026-09-26
- Two local AIs talk to each other with no cloud: hands-on with Braid — Scobleizer · 2026-09-26
- The handiest GPU this dev ever bought is a ~$300 Intel Arc A310, not NVIDIA or AMD — TheZachMueller · 2026-09-26
- Running image generation in the browser on local hardware: 10s pixel art on an RTX 3060 — Bartholomheow · 2026-09-26
- Alibaba's T-Head unveils Zhenwu V900 chip: 216GB per card, sales in Q1 2027 — shashib · 2026-09-26
- llama.cpp fork dedups repeated prompts losslessly, cutting 108k to 71k tokens in agent loops — Odd_Cauliflower_8004 · 2026-09-26