GLM-5.2 local inference: ubatch size significantly boosts MoE performance

fuzhongkai · reddit · 2026-08-22

Tests on GLM-5.2-UD-IQ2XXS GGUF across 3x RTX PRO 6000 Blackwell GPUs reveal that larger ubatch sizes significantly improve throughput for long prompts (763 to 1145 t/s) but not short ones. The author attributes this to the 256-expert / top-8 MoE architecture, where larger batches better utilize expert GEMMs. Additionally, layer splitting outperforms Tensor Parallelism (TP=3) on PCIe without NVLink.

Original post →

More from Infra

Infra channel →