DeepSeek V4.1 Flash weights go live on Hugging Face under MIT license
solyarisoftware · x · 2026-09-10
DeepSeek has released the weights of DeepSeek-V4.1-Flash on Hugging Face under a permissive MIT license.
The repo describes an image-text-to-text multimodal model with text-generation capability, shipped as safetensors with 8-bit and FP8 precision support, loadable directly via transformers. It gathered 184 likes within hours of release.
Full parameter counts and benchmark results have yet to be officially detailed.
Related event: DeepSeek Unveils Open-Source V4.1-Flash MoE Model(28 posts)→
More from Models
- DeepSeek V4.1 Flash is actually 748B params, safetensors analysis shows — DistanceSolar1449 · 2026-09-10
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- Follow-up: a 3T-parameter model may already exist, scaling issues remain the wildcard — teortaxesTex · 2026-09-10
- Speculation: DeepSeek V4.1 Pro could be a 3.1T-param MoE with 2.6TB disk footprint — teortaxesTex · 2026-09-10
- New Book Teaches Beginners to Build and Fine-Tune Their Own GPT-Style SLMs, With Colab Notebooks — Roger_M_Taylor · 2026-09-10