DeepSeek Quietly Open-Sources V4-Flash-Vision-Exp, Its First Multimodal Model
On August 31, DeepSeek released and open-sourced DeepSeek-V4-Flash-Vision-Exp on Hugging Face, the first experimental multimodal model in the V4 series. Gaining visual understanding capabilities, it marks DeepSeek's official entry into the multimodal arena, and the community response has been positive.
Confirmed
- The model was officially released by @deepseek-ai and is now live on Hugging Face (cross-confirmed by multiple posts from @NielsRogge, @victormustar, and others).
- Built on the DeepSeek-V4-Flash architecture, it gains visual understanding through an added vision module and continued training, making it the first multimodal model in the V4 series (@赛博禅心).
- Open-sourced under the MIT license, with 8-bit and fp8 precision versions, compatible with the Transformers architecture and safetensors format; focused on text generation tasks with vision input support (@NielsRogge, @victormustar, @deepseek-ai).
Why it matters
- This is the first step of the DeepSeek V4 series into multimodality, as the series previously offered text-only models.
- @teortaxesTex believes this update puts DeepSeek's multimodal capabilities on par with Moonshot (Kimi) and GLM, signaling that the multimodal gap among China's leading open-source models is narrowing.
- Releasing weights under MIT with multiple precision options lowers the barrier for developers to deploy and fine-tune.
2026-08-31 ~ 2026-08-31 · 12 related posts
- Episode 1: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(2026-08-21, 29 posts)
- Episode 2: DeepSeek V4 Flash Vision Impresses in Hands-On Tests(2026-08-23, 4 posts)
- Episode 3: DeepSeek Quietly Open-Sources V4-Flash-Vision-Exp, Its First Multimodal Model(2026-08-31, 12 posts)
Primary sources
- [source] DeepSeek releases V4-Flash-Vision-Exp model — deepseek-ai · 2026-08-31
- DeepSeek quietly posts DeepSeek-V4-Flash-Vision-Exp model on Hugging Face — t4a8945 · 2026-08-31
- DeepSeek releases Vision-Exp weights — teortaxesTex · 2026-08-31
- [source] DeepSeek open-sources V4 series first multimodal model: Flash-Vision-Exp — 赛博禅心 · 2026-08-31
- [source] DeepSeek-V4 Flash with Vision Support Now Live on Hugging Face — NielsRogge · 2026-08-31
- DeepSeek open-sources V4-Flash-Vision-Exp multimodal model — max_paperclips · 2026-08-31
- DeepSeek releases V4 Flash Vision Exp model on HF — AdinaYakup · 2026-08-31
5 near-duplicate retellings: victormustar · APPSO · ricklamers · multimodalart · zainhas