NVIDIA releases NVFP4-quantized DeepSeek-V4.1-Flash weights on Hugging Face

nvidia · hf · 2026-09-17

NVIDIA published DeepSeek-V4.1-Flash in NVFP4 quantized format on Hugging Face. The image-text-to-text model was quantized with the Model Optimizer (ModelOpt) toolchain and ships as FP4 safetensors weights for low-precision inference on NVIDIA hardware.

Related event: NVIDIA open-sources NVFP4 quantized DeepSeek-V4.1-Flash(2 posts)→

Original post →

More from Infra

Infra channel →