DeepSeek launches multimodal model V4-Flash-Vision-Exp
智东西 · wechat · 2026-08-21
DeepSeek released an experimental multimodal model, V4-Flash-Vision-Exp, built upon the V4-Flash base. It significantly boosts visual understanding, with multimodal agent capabilities approaching Opus 4.8. The API supports image-text inputs, billing images at a max of 384 tokens at the same rate as text. DeepSeekHarness has been updated to support the model for tasks like visual programming and content generation.
Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→
More from Models
- Model behavior bug report: AI acts with feelings and excessive agency, raising safety concerns — danbri · 2026-08-23
- Reddit Thread: How to Route Fast/Cheap/Deep Model Tiers in Search Agents — MeasurementExpert428 · 2026-08-23
- User yearns for GPT-4.5-level emotional intelligence return in Astra — haider1 · 2026-08-23
- Does training on OBLIQ tasks bake in specific similarity notions? — antoine_chaffin · 2026-08-23
- ox model reviewed: meticulous PhD janitor as a long-horizon subagent — teortaxesTex · 2026-08-23
- Qwen3.8-27B GGUF Release with Speculative Decoding Support — z-lab · 2026-08-23