Hugging Face Adds NVFP4 Kernel Support, Cutting VRAM by 61%

ariG23498 · x · 2026-08-21

The article highlights the impact of the NVFP4 quantization format on reducing VRAM usage. For the Muse Glimmer model, NVFP4 cuts VRAM from 59.62 GB to 23.07 GB (a 61% reduction). Crucially, Hugging Face transformers has introduced native NVFP4 GEMM kernel support via PR 47883. This allows models to be computed directly in the NVFP4 format without dequantization, leading to lower memory usage and higher tokens per second.

Original post →

More from Infra

Infra channel →