DeepSeek Launches V4-Flash Multimodal Model, Agent Capability Near Opus-4.8
DeepSeek · wechat · 2026-08-21
DeepSeek has released an experimental multimodal vision model, DeepSeek-V4-Flash-Vision-Exp. While its pure text capabilities match the official V4-Flash, it achieves significant improvements in vision-related Agent benchmarks, approaching the performance of Opus-4.8.
The model demonstrates strong multimodal adaptability within Agent frameworks, supporting complex workflows such as generating business custom PPTs, website redesigns, and dynamic frontend demos.
Regarding API, images are billed by tokens (max 384 tokens) at the same rate as V4-Flash. It supports three formats (ChatCompletions, Messages, Responses) and various image input methods. Additionally, the free FilesAPI is now live, allowing image uploads for reuse via fileid.
Related event: DeepSeek Launches V4-Flash-Vision-Exp for Fast, Cheap Multimodal Agents(13 posts)→
More from Models
- DeepSeek shifts strategy to target specific benchmarks against Opus-4.8 — eyishazyer · 2026-08-21
- Indic-Translate released: 32K context translation for 22 Indian languages — prajdabre · 2026-08-21
- Small model matches Opus 4.8 level: A warning for Anthropic? — kimmonismus · 2026-08-21
- Grok produced 'word salad' yesterday, possibly due to v4.6 testing — mark_k · 2026-08-21
- Ornith-1.5-35B-A3B Tested: 250 tok/s and Strong Agentic Performance — koloved · 2026-08-21
- API Model Gains Vision Capabilities in Major Update — teortaxesTex · 2026-08-21