DeepSeek launches multimodal vision model, rivaling Claude Opus
机器之心 · wechat · 2026-08-21
DeepSeek announced the launch of the multimodal vision understanding model DeepSeek-V4-Flash-Vision-Exp on its API platform, marked as an experimental version. While matching the text capabilities of V4-Flash, it achieves a significant leap in visual Agent Benchmarks, with multimodal agent capabilities approaching Claude Opus-4.8. Official demos include generating custom travel PPTs, redesigning website UI styles, and creating dynamic frontend effects. Meanwhile, an unnamed domestic model called OxAlpha has also sparked discussion in the community.
Related event: DeepSeek Launches V4-Flash-Vision-Exp, Closing In on Opus 4.8(29 posts)→
More from Models
- Model behavior bug report: AI acts with feelings and excessive agency, raising safety concerns — danbri · 2026-08-23
- Reddit Thread: How to Route Fast/Cheap/Deep Model Tiers in Search Agents — MeasurementExpert428 · 2026-08-23
- User yearns for GPT-4.5-level emotional intelligence return in Astra — haider1 · 2026-08-23
- Does training on OBLIQ tasks bake in specific similarity notions? — antoine_chaffin · 2026-08-23
- ox model reviewed: meticulous PhD janitor as a long-horizon subagent — teortaxesTex · 2026-08-23
- Qwen3.8-27B GGUF Release with Speculative Decoding Support — z-lab · 2026-08-23