DeepSeek Launches Multimodal V4-Flash-Vision Model for Agents with Files API

APPSO · wechat · 2026-08-21

DeepSeek released the experimental deepseek-v4-flash-vision-exp model, adding vision capabilities to the V4-Flash text base to enhance Agent task performance. It performs close to high-end models like Opus-4.8 in multimodal Agent benchmarks. Pricing follows V4-Flash, with images charged by tokens (up to 384 tokens/image). The accompanying Files API allows uploading JPEG/PNG files and referencing them via fileid across requests, reducing redundancy for complex workflows requiring repeated image analysis.

Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→

Original post →

More from coding & agent

coding & agent channel →