DeepSeek opens its eyes: V4-Flash-Vision lands at ~9 images per cent
量子位 · wechat · 2026-08-22
After a four-month wait, DeepSeek released deepseek-v4-flash-vision-exp, a multimodal experimental model with native image input (no image generation). It's live on API at Flash pricing: one image costs at most 384 tokens — roughly ¥0.0012 per image under peak/no-cache rates, about 9 images per cent, far cheaper than a competitor's ¥0.02.
QbitAI's hands-on: text capabilities (agent, reasoning, world knowledge) match V4 Flash, but on multimodal agent benchmarks it jumps significantly, approaching Opus-4.8; basic image recognition feels identical to the beta web model. Paired with DeepSeekHarness (updated the same day with native first-party multimodal support), it handled image-to-3D-modeling in Blender and Canvas front-end demos well — though finger-counting remains a common failure mode. The post also mentions the viral "牛来" model on X, widely speculated to be from Zhipu.
Related event: DeepSeek Releases V4-Flash-Vision-Exp Multimodal Model(26 posts)→
More from Models
- NVIDIA's AVO harness lifts Opus 5 from 30% to 100% on ARC-AGI-3 — daniel_mac8 · 2026-08-22
- Qwen 3.8 vs 3.6: Low reasoning mode loops less — Lair98 · 2026-08-22
- GLM-5.3 Takes 2nd Place on Creative Writing Benchmark with Qualitative Analysis — zero0_one1 · 2026-08-22
- Private Benchmark: 0x Alpha Underperforms on Low Reasoning Tasks — toptickcrypto · 2026-08-22
- Ox Alpha Model Test: Surprising Performance in Code and Data Analysis — killerbee1432 · 2026-08-22
- Qwen3.8-27B: 3x Faster Long-Context Decoding with DFlash2 + XQA — stargate425 · 2026-08-22