NVIDIA releases NVFP4-quantized DeepSeek-V4.1-Flash weights on Hugging Face
nvidia · hf · 2026-09-17
NVIDIA published DeepSeek-V4.1-Flash in NVFP4 quantized format on Hugging Face. The image-text-to-text model was quantized with the Model Optimizer (ModelOpt) toolchain and ships as FP4 safetensors weights for low-precision inference on NVIDIA hardware.
Related event: NVIDIA open-sources NVFP4 quantized DeepSeek-V4.1-Flash(2 posts)→
More from Infra
- Jensen Huang: a 1GW NVIDIA AI factory costs $50-60B but generates ~$50B in annual rental revenue — rohanpaul_ai · 2026-09-17
- Ilya warns neoclouds' weak cybersecurity invites rogue AI agents to hijack compute — Miles_Brundage · 2026-09-17
- AMD's free AI Developer Program: $100 cloud credits, Discord access, hardware raffles — wkmyrhang · 2026-09-17
- After AWS me-central-1 loss, dev jokes about explaining the outage to Codex weekly — andersonbcdefg · 2026-09-17
- Crusoe runs 512 AMD MI355X GPUs at 5.75M tok/s in largest MLPerf inference entry — wkmyrhang · 2026-09-17
- Perovskite could lift solar efficiency ceiling from 30% to 45% — and give the US a shot against China — kyliebytes · 2026-09-17