DeepSeek launches multimodal model V4-Flash-Vision-Exp

智东西 · wechat · 2026-08-21

DeepSeek released an experimental multimodal model, V4-Flash-Vision-Exp, built upon the V4-Flash base. It significantly boosts visual understanding, with multimodal agent capabilities approaching Opus 4.8. The API supports image-text inputs, billing images at a max of 384 tokens at the same rate as text. DeepSeekHarness has been updated to support the model for tasks like visual programming and content generation.

Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→

Original post →

More from Models

Models channel →