NVIDIA releases NVFP4 quantized DeepSeek-V4.1-Flash on Hugging Face
TheZachMueller · x · 2026-09-17
NVIDIA has published an NVFP4 quantized build of DeepSeek-V4.1-Flash on Hugging Face (repo: nvidia/DeepSeek-V4.1-Flash-NVFP4), produced with its Model Optimizer toolchain.
- The model is an image-text-to-text multimodal model based on deepseek-ai/DeepSeek-V4.1-Flash
- Uses FP4/8-bit quantization for lower memory footprint and faster inference
- Released under the MIT license, remaining open source
For developers running DeepSeek V4.1 locally or on constrained hardware, this offers a ready-made low-precision deployment option.
Related event: NVIDIA open-sources NVFP4 quantized DeepSeek-V4.1-Flash(2 posts)→
More from Infra
- llama.cpp fails to load Qwen3.8 MTP draft model: 'output_hc_norm.weight' tensor not found — Ambitious_Fold_2874 · 2026-09-17
- Prediction: Kimi K3-level AI on a single RTX 5090 within 18 months — TheZachMueller · 2026-09-17
- User says he'd pay $1k/month for AI, but high pricing makes local AI attractive — draginol · 2026-09-17
- Tencent open-sources FlexKV distributed KV cache for LLM inference, cutting TTFT by up to 70% — Roger_M_Taylor · 2026-09-17
- Optimization mined via Bittensor competition lands in vLLM, boosting Qwen3 throughput ~4% — const_reborn · 2026-09-17
- Early vLLM PR adds Jev-like structured generation for DiffusionGemma, only 2x endpoint latency on a DGX Spark — generativist · 2026-09-17