GLM-5.2 NVFP4 Inference Speedup

nickcraske · x · 2026-07-10

A cited source reveals that GLM-5.2 has achieved inference optimization on Nvidia NVFP4 without pruning or quantization, retaining the same model card and delivering 110 tok/s in a single stream. The quote also mentions a repo reproducing Full DS4-Flash running at 38.5 tok/s on a 5090 + DDR5 setup, which the author calls a breakthrough.

Original post →

More from Infra

Infra channel →