DeepSeek releases V4 Flash Vision Exp model on HF

AdinaYakup · x · 2026-08-31

DeepSeek released a new model, DeepSeek-V4-Flash-Vision-Exp, on Hugging Face. It is a multimodal model (image-text-to-text) capable of processing visual and text inputs. It scores 83.9 on Terminal Bench 2.1 (Rank 8) and 59.3 on Deep SWE (Rank 5). The model supports Transformers, Safetensors, 8-bit, and FP8 quantization.

Related event: DeepSeek Quietly Open-Sources V4-Flash-Vision-Exp, Its First Multimodal Model(12 posts)→

Original post →

More from Multimodal

Multimodal channel →