DeepSeek open-sources 305B V4-Flash-Vision-Exp, an experimental vision model built for agents
大模型之路 · wechat · 2026-09-02
DeepSeek released DeepSeek-V4-Flash-Vision-Exp on HuggingFace, the first experimental multimodal model in its V4 family: 305B total parameters, only 13B active, MIT-licensed. Unlike captioning models, it is explicitly positioned for multimodal agents — reading screenshots, UIs and charts directly, then completing tasks via tool calls.
Benchmarks show vision barely hurts text performance: TerminalBench 2.1 rose from 82.7 to 83.9, DeepSWE from 54.4 to 59.3; Agents' LastExam hit 27.3, beating Opus 4.8's 25.7. DeepSeek says agent capabilities are “close to Claude Opus 4.8”. Community quantizations for llama.cpp, Ollama and LMStudio are already out, and the API offers the vision features for quick testing before local deployment.
Caveats: the 305B size suits server-side agents, not edge devices; fine-grained UI grounding and long-horizon state tracking remain weak spots, and benchmark gains may not transfer to your own UIs — test with real screenshots first.
More from coding & agent
- Giving AI agents their own inbox is architecturally wrong, Reddit thread argues — Creamy-And-Crowded · 2026-09-02
- Merge launches enterprise AI governance tool enforcing model routing and spend rules — shensi · 2026-09-02
- Weaviate Ask Mode Adds Configurable Evaluation to Trade Latency vs Verifiability — CShorten30 · 2026-09-02
- Building the Ultimate Agent Harness for Kimi K3: The Model Is No Longer the Bottleneck — VibeMarketer_ · 2026-09-02
- GLM 5.2 slug references spotted in Google Antigravity CLI, hinting at integration — gaganghotra_ · 2026-09-02
- Design lead ships 12 PRs in a week: AI is erasing the designer-engineer gap — talkaboutdesign · 2026-09-02