Paired 4:8 sparsity gives 1.35-1.65x over dense NVFP4 on B200, 1.18x in serving
vllm_project · x · 2026-10-07
The vLLM project highlights @jbish1572's data on paired 4:8 sparsity, where each group of eight weights keeps two adjacent pairs. Expert GEMMs ran 1.35-1.65x faster than dense NVFP4 on B200, but the theoretical 2x peak doesn't survive real serving: Kimi-K2.5 in vLLM achieved 1.18x. A useful reality check for sparse MoE inference gains.
More from Infra
- Nvidia nears $6T market cap with record $150B buyback; SpaceX in talks to borrow $40B for chips — rohanpaul_ai · 2026-10-07
- 7 minutes per motor, ~$5 labor cost: why robotics automation is the only path for Western manufacturing — avlok · 2026-10-07
- Railroad buildout ran at 1.5-2% of GDP for 60 years — AI capex only hit that level this year — toptickcrypto · 2026-10-07
- Same GPUs, 100x Gap: vLLM Hits 110 tok/s on Dual RTX 5090s Where llama.cpp Took 30 Minutes — vllm_project · 2026-10-07
- kipply's July-August digest: export controls lifted, alignment-faking paper, and more — kipperrii · 2026-10-07
- Marvell Investor Day 2026 materials land, with a nudge to fix the chart arrows — jwt0625 · 2026-10-07