DeepSeek V4.1-Flash: KV cache 437x smaller, ~40x cheaper than Claude Opus 4.8
AdinaYakup · x · 2026-09-10
Analysis of DeepSeek V4.1 Flash shows striking specs:
- Asymmetric causal encoder-decoder: 550B MoE with 8B active on input / 16B on output
- Native vision merged into a single endpoint
- KV cache crushed: 1/4 the HBM of the last gen, 437x smaller than their first model
- Pricing: $0.6/M output off-peak — 6x cheaper than flagship Flash models, 40x under Claude Opus 4.8 and GPT-5.6
- The "Flash" variant even beats their own flagship V4 Pro
The model is open on Hugging Face under MIT license.
Related event: Leaked DeepSeek V4.1 Benchmarks Point to New 552B Architecture(15 posts)→
More from Infra
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- After Nvidia's Hugging Face buyout, devs call for a neutral alternative — hargup13 · 2026-09-10
- Google Cloud user hit with an $82k bill within 5 hours — Patient_Election2179 · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Dual RTX Pro 6000 + Threadripper 9955W local LLM build — sanity check requested — No_Run8812 · 2026-09-10