DeepSeek Releases V4 Flash Vision API for Image Understanding
aigclink · x · 2026-08-21
DeepSeek has released API documentation for the deepseek-v4-flash-vision-exp model, enabling image understanding within text conversations. The model supports describing images, OCR, and analyzing charts.
Formats & Limits:
- Supports JPEG, PNG, GIF, and WebP.
- Three input methods:
- Base64 Inline: Best for local files, subject to a 48MiB request limit.
- Public URL: Max 32MiB, must download within 60s.
- Files API: Best for reuse or large files (up to 64MiB), referenced via fileid.
The API is OpenAI-compatible, with Python and cURL examples provided.
Related event: DeepSeek Launches Experimental Multimodal Model V4-Flash-Vision-Exp(18 posts)→
More from Models
- Ox Alpha Coming Soon, Excitement Builds — zephyr_z9 · 2026-08-21
- Ox-Alpha's take on researcher Janus: real contributions, polarizing figure — jd_pressman · 2026-08-21
- mLateOn matches 8B models on Japanese JMTEB, beats Japan-specific models — IgorCarron · 2026-08-21
- User Report: Opus 5 Regresses, Context Compaction Losing Track — springrod · 2026-08-21
- The Forecasting Company open-sources t0-alpha, detailing the 5-era evolution of forecasting models — fpedregosa · 2026-08-21
- SentenceTransformers V6 Unifies Text, Vision, and Audio in Single API — ManuelFaysse · 2026-08-21