DeepSeek open-sources V4 series first multimodal model: Flash-Vision-Exp
赛博禅心 · wechat · 2026-08-31
DeepSeek has open-sourced DeepSeek-V4-Flash-Vision-Exp, the first experimental multimodal model in the V4 series. Built on the DeepSeek-V4-Flash architecture, it gains visual understanding capabilities through introduced vision modules and continuous training. Compared to V4-Flash-0731, it shows significant improvements in multimodal Agent capabilities while maintaining comparable performance in pure text Agent tasks.
More from Models
- Insider: Model Evals Interrupted Because Compute Was Reclaimed for Other Evaluations — sjgadler · 2026-08-31
- Developer Switches from Opus 5 to Codex Citing Better Performance — sirbayes · 2026-08-31
- Dots3-Note Weights Open: Addressing the Benchmark-Reality Gap — CodeByPoonam · 2026-08-31
- Tiel-Coder-35B-A3B Trends on HF with Speculative Decoding & MTP — peculiar-ragdoll · 2026-08-31
- Google Releases Gemini Omni 1.1 Flash, Updating Its Fast Multimodal Model for Developers — thione · 2026-08-31
- DeepSeek launches low-cost vision model; Anthropic previews hardware control protocol for agents — thione · 2026-08-31