Economics of Deploying GLM-5.2 on 8×B200
qubridInc · reddit · 2026-07-08
The Qubrid team shared a deployment configuration analysis for GLM-5.2 (approx. 750B total params/40B activated MoE, 256 experts top-8, DSA+MLA attention, 1M context, MIT license) on 8×B200 nodes. Key findings: MoE decoding is bandwidth-bound rather than compute-bound at medium concurrency, meaning B200 offers only about 1.2x performance per dollar compared to H200 at FP8 (matching the HBM bandwidth ratio, not the 2.3x compute ratio). The real game-changer is NVFP4 (halving weight bytes, which Hopper lacks FP4 tensor cores for). The recommended configuration is NVFP4 + two TP=4 replicas, achieving roughly 2x throughput compared to TP=8. The author also noted the current lack of a comprehensive concurrency sweep table for GLM-5.2 on a documented 8×B200 setup.
More from Infra
- WSJ: Nvidia is in talks to backstop about $250 billion of OpenAI's data center plan — KateClarkTweets · 2026-07-27
- YC talk on BCI x AI says infrastructure is what really determines speed — garrytan · 2026-07-27
- A 13B model ran on a no-GPU PC by paging weights from SSD via llama.cpp — ID_R_McGregor · 2026-07-27
- llama.cpp warns that GGUFs made before a recent change must be regenerated — EconomySerious · 2026-07-27
- RTX 5090 local tests show Qwen Q6 can drop to 15 tok/s at 80k context — LFAdvice7984 · 2026-07-27
- Surprising Ubuntu Setup: NVIDIA 5090 PC Becomes the Easiest AI Rig — _xjdr · 2026-07-27