DeepSeek launches multimodal API with 384 tokens per image, 25x cheaper than rivals
新智元 · wechat · 2026-08-21
DeepSeek has released its first multimodal API model, DeepSeek-V4-Flash-Vision-Exp, featuring strong vision capabilities that outperform predecessors in six out of seven Agent benchmarks. The model supports standard API formats and is integrated with the DeepSeek Harness framework. Crucially, it prices images at a maximum of 384 tokens, identical to its text-only model. This results in a cost of approximately 1.15 RMB for 1,000 images (off-peak), which is 25x to 50x cheaper than competitors like Claude and GPT-4o. Tests demonstrate the model's ability to generate 1:1 HTML code from screenshots and create professional B2B pitch decks from vague prompts.
Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→
More from Multimodal
- Rough tier list of every major vision model — andrejusb · 2026-08-23
- Saxxy Harry video made with LTX 2.5, Suno, and ElevenLabs — Kyrannio · 2026-08-23
- Funny dance video with twist ending made with Seedance 2.5 — Kyrannio · 2026-08-23
- AI-generated video stuns with realism, sparking model speculation — Xianbao_QIAN · 2026-08-23
- A Reusable Product Photo Prompt Template for Sculptural Editorial Scenes — azed_ai · 2026-08-23
- VideoCoCo Fixes Video Physics by Generating Executable Code — TinfoilTricorn · 2026-08-23