DeepSeek releases experimental V4 Flash Vision multimodal model

victormustar · x · 2026-08-31

DeepSeek has released a new experimental model, DeepSeek-V4-Flash-Vision-Exp, on Hugging Face. This multimodal model is based on the Transformer architecture and focuses on text-generation tasks with vision capabilities. It is licensed under MIT, available in 8-bit and fp8 precisions, and is compatible with the Transformers library and Inference Endpoints.

Related event: DeepSeek Quietly Open-Sources V4-Flash-Vision-Exp, Its First Multimodal Model(12 posts)→

Original post →

More from Multimodal

Multimodal channel →