A GLM 5.2 serving rig reaches 47 tok/s/user after async pipeline tuning

TheZachMueller · x · 2026-07-21

A cursed GLM 5.2 rig hits 47 tok/s/user after async pipeline tuning

The author describes a highly customized serving setup for GLM 5.2 NVFP4 that ran through 6.4M tokens over 10h47m of optimization.

The attached chart shows the step-by-step gains from baseline to final peak.

Related event: Extreme Topology Tuning Boosts LLM Inference Throughput to 47 tok/s(2 posts)→

Original post →

More from Infra

Infra channel →