DeepSeek releases V4 Flash Vision Exp model on HF
AdinaYakup · x · 2026-08-31
DeepSeek released a new model, DeepSeek-V4-Flash-Vision-Exp, on Hugging Face. It is a multimodal model (image-text-to-text) capable of processing visual and text inputs. It scores 83.9 on Terminal Bench 2.1 (Rank 8) and 59.3 on Deep SWE (Rank 5). The model supports Transformers, Safetensors, 8-bit, and FP8 quantization.
More from Multimodal
- Seedance 2.5 Generates Photorealistic IGI-Style Tactical FPS Gameplay — SimplyAnnisa · 2026-08-31
- Lovart launches China edition: one agent delivers full brand design kits, reference plugin and 100 pro skills — 卡尔的AI沃茨 · 2026-08-31
- ChatGPT Images Can Now Edit Photos with Prompts, Rivaling Photoshop — TawohAwa · 2026-08-31
- Can MiniMax H3 run on 8GB VRAM and 16GB RAM or is it pointless? — Shanq123 · 2026-08-31
- DeepSeek-V4-Flash-Vision-Exp Open Sourced, Topping Multimodal Benchmarks — 智东西 · 2026-08-31
- Fixing AI video face melt: upload raw audio and quote dialogue in prompt — Fragrant-Cheek-4273 · 2026-08-31