B200 benchmark: NVFP4 decodes up to ~8% faster than MXFP4 in low-batch LLM inference
StasBekman · x · 2026-10-11
Cezar Cocu published a production-workload benchmark comparing NVFP4 and MXFP4, going beyond format-level comparisons.
- Setup: Qwen3-32B dense model served with vLLM on a B200, both formats quantized from scratch from the same BF16 weights to minimize confounders.
- Verdict: On B200, prefer NVFP4 across the board — up to 8% faster decode at small batches, and it also looks better for GEMM-bound training (9% more TFLOPS on large GEMMs, kernel-dependent).
- Counterintuitive finding: the gap disappears at larger batch sizes; the author initially assumed HBM bandwidth saturation but analysis showed that wasn't the reason.
- The benchmark was mostly built by Claude and run by the author; methodology is reproducible.
More from Infra
- Running GLM-5.3-Flash on dual Ascend 310P cards: 8-9 tok/s and 311K context — matteiuspi · 2026-10-12
- Anthropic subscribers reportedly get 4X more compute per dollar than API buyers as margins hit 88% — rohanpaul_ai · 2026-10-12
- Can killing NDAs win over data center opponents? Equity podcast weighs in — TechCrunch AI · 2026-10-12
- New GPU benchmark data: B300 hits 84% of theoretical, B200 only 77% — StasBekman · 2026-10-12
- Fractile AI CEO on why third-party chips survive the 5x efficiency race — Tom_Westgarth15 · 2026-10-12
- Dev cuts 50ms per layer in GPU+CPU hybrid local LLM inference with hot-expert mapping — HankYeomans · 2026-10-12