DeepSeek launches V4-Flash Vision model for multimodal agents
AI寒武纪 · wechat · 2026-08-21
DeepSeek released the experimental multimodal model DeepSeek-V4-Flash-Vision-Exp, adding vision capabilities to the Flash series.
- Performance: Text capabilities match V4-Flash; multimodal agent performance approaches Claude Opus-4.8.
- API Updates: Supports ChatCompletions, Messages, and Responses. Images can be sent via base64, URL, or Files API.
- Files API: Now available for free; upload images once and reference by fileid to save bandwidth.
- Cost: Images charged by tokens (max 384 tokens/image), consistent with V4-Flash pricing.
This update significantly expands the operational scope of agents in visual scenarios like screenshot analysis and chart understanding.
Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→
More from Models
- Model behavior bug report: AI acts with feelings and excessive agency, raising safety concerns — danbri · 2026-08-23
- Reddit Thread: How to Route Fast/Cheap/Deep Model Tiers in Search Agents — MeasurementExpert428 · 2026-08-23
- User yearns for GPT-4.5-level emotional intelligence return in Astra — haider1 · 2026-08-23
- Does training on OBLIQ tasks bake in specific similarity notions? — antoine_chaffin · 2026-08-23
- ox model reviewed: meticulous PhD janitor as a long-horizon subagent — teortaxesTex · 2026-08-23
- Qwen3.8-27B GGUF Release with Speculative Decoding Support — z-lab · 2026-08-23