Economics of Deploying GLM-5.2 on 8×B200
qubridInc · reddit · 2026-07-08
The Qubrid team shared a deployment configuration analysis for GLM-5.2 (approx. 750B total params/40B activated MoE, 256 experts top-8, DSA+MLA attention, 1M context, MIT license) on 8×B200 nodes. Key findings: MoE decoding is bandwidth-bound rather than compute-bound at medium concurrency, meaning B200 offers only about 1.2x performance per dollar compared to H200 at FP8 (matching the HBM bandwidth ratio, not the 2.3x compute ratio). The real game-changer is NVFP4 (halving weight bytes, which Hopper lacks FP4 tensor cores for). The recommended configuration is NVFP4 + two TP=4 replicas, achieving roughly 2x throughput compared to TP=8. The author also noted the current lack of a comprehensive concurrency sweep table for GLM-5.2 on a documented 8×B200 setup.
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11