DeepSeek open-sources V4's first multimodal model, targeting Agent vision capabilities

APPSO · wechat · 2026-08-31

DeepSeek has officially open-sourced DeepSeek-V4-Flash-Vision-Exp, the first experimental multimodal model in the V4 family, under the MIT license. Built upon the DeepSeek-V4-Flash architecture with added vision modules, it retains text, reasoning, and Agent capabilities while gaining image understanding. The release includes model files, Tokenizer, and a minimal PyTorch inference implementation.

The model is positioned towards Agent use cases, emphasizing Multimodal Agent Capabilities to read visual information like web screenshots and software interfaces for tool execution. Benchmarks show that with vision capabilities, TerminalBench 2.1 rose to 83.9 and DeepSWE to 59.3, outperforming Opus-4.8. In multimodal Agent tests, ApexBench Pass@1 reached 36.5, also surpassing the baseline.

Related event: DeepSeek Quietly Open-Sources V4-Flash-Vision-Exp, Its First Multimodal Model(12 posts)→

Original post →

More from coding & agent

coding & agent channel →