A GLM 5.2 serving rig reaches 47 tok/s/user after async pipeline tuning
TheZachMueller · x · 2026-07-21
A cursed GLM 5.2 rig hits 47 tok/s/user after async pipeline tuning
The author describes a highly customized serving setup for GLM 5.2 NVFP4 that ran through 6.4M tokens over 10h47m of optimization.
- Starting point: a conventional 6-GPU multi-node vLLM baseline at about 20 tok/s/user.
- They moved to an asymmetric TP2→TP4 topology with a minimal inter-node pipeline boundary.
- A 30/48 layer split plus async PP/NCCL work improved throughput while preserving exact outputs.
- Final result: 47.1 tok/s/user at T=1 and 42.7 tok/s/user at T=4, with 169 tok/s aggregate at four users.
The attached chart shows the step-by-step gains from baseline to final peak.
Related event: Extreme Topology Tuning Boosts LLM Inference Throughput to 47 tok/s(2 posts)→
More from Infra
- NVIDIA details Vera CPU with 2x performance claims and a 22,000-core rack — ryanshrout · 2026-07-22
- NVIDIA says Vera Rubin NVL72 delivers 10x more tokens per megawatt than Blackwell — nvidia · 2026-07-22
- AI economy is running out of cheap compute as data-center and power costs surge — Scobleizer · 2026-07-22
- Ratel says it made agents 7x cheaper by loading only the tools each task needs — tensorqt · 2026-07-22
- Weaviate adds per-query profiling to pinpoint where a slow search query spends time — CShorten30 · 2026-07-22
- Mistral expands its Microsoft partnership as it adds more AI compute in Europe — MistralAI · 2026-07-22