DeepSeek Releases V4.1-Flash with 1M Token Context

DeepSeek released V4.1-Flash, its smallest new-architecture model, featuring a 1M token context, 4x smaller KV cache, and a multimodal MoE design with 552B backbone parameters (763B total). It is now available on the FLock API platform and Ollama's cloud mode.

2026-09-11 ~ 2026-09-11 · 2 related posts

Full story(16 episodes)→