GLM 5.2 Hits 385 tok/s on 8x RTX PRO 6000 with NVFP4

max_paperclips · x · 2026-08-08

A developer successfully ran the GLM 5.2 model on 8x RTX PRO 6000 GPUs using the NVFP4 quantization format, achieving an inference speed of 385 tokens/s.

The poster noted that initial runs were buggy and incredibly slow, taking over a week of debugging with the help of AI agents to reach this performance milestone.

Original post →

More from Infra

Infra channel →