DeepSeek-V4.1-Flash hits Ollama: 552B MoE backbone with 1M context via KV cache compression

ollama · x · 2026-09-11

DeepSeek-V4.1-Flash is now available on Ollama (cloud mode). The multimodal MoE model has a 552B backbone and supports up to 1M-token contexts, with vision, tools, and thinking modes.

Original post →

More from Infra

Infra channel →