DeepSeek Open-Sources V4 Flash Vision Multimodal Model

ccerrato147 · x · 2026-09-01

DeepSeek has released the DeepSeek-V4-Flash-Vision-Exp model as open-weight. This is an image-text-to-text multimodal model based on the Transformer architecture. It achieved a score of 83.9 (rank 8) on Terminal Bench 2.1 and 59.3 (rank 5) on Deep Swe. The model supports 8-bit and fp8 quantization and is licensed under MIT.

Related event: DeepSeek quietly open-sources V4-Flash-Vision-Exp, first multimodal model in V4 family(13 posts)→

Original post →

More from Multimodal

Multimodal channel →