DeepSeek launches multimodal API with 384 tokens per image, 25x cheaper than rivals

新智元 · wechat · 2026-08-21

DeepSeek has released its first multimodal API model, DeepSeek-V4-Flash-Vision-Exp, featuring strong vision capabilities that outperform predecessors in six out of seven Agent benchmarks. The model supports standard API formats and is integrated with the DeepSeek Harness framework. Crucially, it prices images at a maximum of 384 tokens, identical to its text-only model. This results in a cost of approximately 1.15 RMB for 1,000 images (off-peak), which is 25x to 50x cheaper than competitors like Claude and GPT-4o. Tests demonstrate the model's ability to generate 1:1 HTML code from screenshots and create professional B2B pitch decks from vague prompts.

Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→

Original post →

More from Multimodal

Multimodal channel →