DeepSeek releases V4.1-Flash: multimodal MoE with 1M token context

DeepSeek launched V4.1-Flash, its smallest new-architecture model: a 552B-parameter multimodal MoE with native vision, 1M-token context, 4x smaller KV cache, and claimed setup costs 1/40 of Claude's. It is now available on Fireworks, FLock, and Ollama.

2026-09-11 ~ 2026-09-12 · 4 related posts

Full story(18 episodes)→