DeepSeek Launches Multimodal V4-Flash-Vision Model for Agents with Files API
APPSO · wechat · 2026-08-21
DeepSeek released the experimental deepseek-v4-flash-vision-exp model, adding vision capabilities to the V4-Flash text base to enhance Agent task performance. It performs close to high-end models like Opus-4.8 in multimodal Agent benchmarks. Pricing follows V4-Flash, with images charged by tokens (up to 384 tokens/image). The accompanying Files API allows uploading JPEG/PNG files and referencing them via fileid across requests, reducing redundancy for complex workflows requiring repeated image analysis.
Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→
More from coding & agent
- Open source MCP server implements HTTP 402 micropayment gateway — EstablishmentTough18 · 2026-08-23
- nb.nvim Fixes Cross-Folder Image Link Bug — 4310sy · 2026-08-23
- AI Agent Search Optimization: Path Calculation Drops from 504ms to Near 0ms — DanielLockyer · 2026-08-23
- How To Build An AI Agent: A Complete Workflow Guide — mdancho84 · 2026-08-23
- A practical guide to AI agents for small businesses: lead response in minutes, 40 hours of invoice work down to 6 — juliesweetie93 · 2026-08-23
- GSD Pi: A terminal agent for long-haul autonomous software projects — tom_doerr · 2026-08-23