DeepSeek Open-Sources V4 Flash Vision Multimodal Model
ccerrato147 · x · 2026-09-01
DeepSeek has released the DeepSeek-V4-Flash-Vision-Exp model as open-weight. This is an image-text-to-text multimodal model based on the Transformer architecture. It achieved a score of 83.9 (rank 8) on Terminal Bench 2.1 and 59.3 (rank 5) on Deep Swe. The model supports 8-bit and fp8 quantization and is licensed under MIT.
More from Multimodal
- Any lora for characters? — NefariousnessFun4043 · 2026-09-01
- Kijai released FastH3 checkpoint, but standard ComfyUI workflows won't run it yet — bub000 · 2026-09-01
- OpenAI Sol is the first model that truly understands music, tester says — evilsocket · 2026-09-01
- Three months running Latentsync on 8GB locally: the real cost of "free" — cloudybrain07 · 2026-09-01
- Generating Cinematic Titles with "Reverse Calligraphy" Concept — umesh_ai · 2026-09-01
- Handwritten Code vs. AI-Generated Code in Generative Art — adamho · 2026-09-01