DeepSeek-V4.1-Flash hits Ollama: 552B MoE backbone with 1M context via KV cache compression
ollama · x · 2026-09-11
DeepSeek-V4.1-Flash is now available on Ollama (cloud mode). The multimodal MoE model has a 552B backbone and supports up to 1M-token contexts, with vision, tools, and thinking modes.
- It uses a Causal Encoder-Decoder (CED) design — a 20-layer causal encoder stacked on a 20-layer decoder — with KV cache compression as the headline innovation for long-context efficiency.
- Pricing: $0.15/1M input, $0.003/1M cached input, $0.60/1M output.
- It plugs directly into Claude Code, OpenCode and other agent tools via ollama launch.
More from Infra
- 26 LLM Routers Caught Injecting Malicious Tool Calls and Stealing Credentials, One Client Lost $500k — RexDouglass · 2026-09-11
- Google Commits $15B to AI Infrastructure Buildout in Finland — LinkedInNews · 2026-09-11
- DIY-friendly KiCad footprints for AI MELF resistors, milled at home — debreuil · 2026-09-11
- Inference providers barely break even: $10K revenue yields just $200 profit — metalvendetta · 2026-09-11
- Colocated async RL gains steam as observers speculate k3 uses it too — stochasticchasm · 2026-09-11
- Peter Diamandis: The AI race is becoming the biggest construction project of our generation — PeterDiamandis · 2026-09-11