GLM-5.2 local inference: ubatch size significantly boosts MoE performance
fuzhongkai · reddit · 2026-08-22
Tests on GLM-5.2-UD-IQ2XXS GGUF across 3x RTX PRO 6000 Blackwell GPUs reveal that larger ubatch sizes significantly improve throughput for long prompts (763 to 1145 t/s) but not short ones. The author attributes this to the 256-expert / top-8 MoE architecture, where larger batches better utilize expert GEMMs. Additionally, layer splitting outperforms Tensor Parallelism (TP=3) on PCIe without NVLink.
More from Infra
- Ox Alpha's 100T tokens/day giveaway costs $1-2M in electricity at full capacity — teortaxesTex · 2026-08-22
- Seeking unified AI gateway for OpenAI cost visibility — HurryOrganic · 2026-08-22
- Agents fail silently: $47k loop reveals monitoring gaps — alifgokce · 2026-08-22
- Anti-datacenter movement is humanity recognizing its successor — ZeroStateReflex · 2026-08-22
- Can llama.cpp share KV cache across multiple GPUs for parallel requests? — spaceman_ · 2026-08-22
- FreeToken: 4x Faster Decode, Enables 284B Models on Gaming Desktops — airesearch12 · 2026-08-22