DeepSeek opens its eyes: V4-Flash-Vision lands at ~9 images per cent

量子位 · wechat · 2026-08-22

After a four-month wait, DeepSeek released deepseek-v4-flash-vision-exp, a multimodal experimental model with native image input (no image generation). It's live on API at Flash pricing: one image costs at most 384 tokens — roughly ¥0.0012 per image under peak/no-cache rates, about 9 images per cent, far cheaper than a competitor's ¥0.02.

QbitAI's hands-on: text capabilities (agent, reasoning, world knowledge) match V4 Flash, but on multimodal agent benchmarks it jumps significantly, approaching Opus-4.8; basic image recognition feels identical to the beta web model. Paired with DeepSeekHarness (updated the same day with native first-party multimodal support), it handled image-to-3D-modeling in Blender and Canvas front-end demos well — though finger-counting remains a common failure mode. The post also mentions the viral "牛来" model on X, widely speculated to be from Zhipu.

Related event: DeepSeek Releases V4-Flash-Vision-Exp Multimodal Model(26 posts)→

Original post →

More from Models

Models channel →