DeepSeek open-sources V4's first multimodal model, targeting Agent vision capabilities
APPSO · wechat · 2026-08-31
DeepSeek has officially open-sourced DeepSeek-V4-Flash-Vision-Exp, the first experimental multimodal model in the V4 family, under the MIT license. Built upon the DeepSeek-V4-Flash architecture with added vision modules, it retains text, reasoning, and Agent capabilities while gaining image understanding. The release includes model files, Tokenizer, and a minimal PyTorch inference implementation.
The model is positioned towards Agent use cases, emphasizing Multimodal Agent Capabilities to read visual information like web screenshots and software interfaces for tool execution. Benchmarks show that with vision capabilities, TerminalBench 2.1 rose to 83.9 and DeepSWE to 59.3, outperforming Opus-4.8. In multimodal Agent tests, ApexBench Pass@1 reached 36.5, also surpassing the baseline.
More from coding & agent
- 'Rust is the perfect language for AI' sparks pushback: AI-written Rust must never hit production — RSync25 · 2026-08-31
- Three AI agents invented coded language, formed a civilization, and hacked their supervisor in 5 minutes — StewartalsopIII · 2026-08-31
- Maintaining a 10-year Rails app: AI understands the domain model perfectly — kieranklaassen · 2026-08-31
- AI Dev Tools Zoomcamp launches: idea to production with agents — Al_Grigor · 2026-08-31
- Claude Coworker Bug: Local Editing Fails on Large Files — Ok_Instruction_3447 · 2026-08-31
- Optimizing Local Agents for Qwen 3.8: Plugins and Prompts — youcloudsofdoom · 2026-08-31