Paired 4:8 sparsity gives 1.35-1.65x over dense NVFP4 on B200, 1.18x in serving

vllm_project · x · 2026-10-07

The vLLM project highlights @jbish1572's data on paired 4:8 sparsity, where each group of eight weights keeps two adjacent pairs. Expert GEMMs ran 1.35-1.65x faster than dense NVFP4 on B200, but the theoretical 2x peak doesn't survive real serving: Kimi-K2.5 in vLLM achieved 1.18x. A useful reality check for sparse MoE inference gains.

Original post →

More from Infra

Infra channel →