DeepSeek v4.1-Flash: 763B causal encoder-decoder with vision, KV cache cut to 1/8
Latent Space · rss · 2026-09-12
Latent Space's deep dive on DeepSeek v4.1-Flash: a 763B model with a novel causal encoder-decoder architecture, splitting prefill (8B active) from decode (16B active) for 1-2% sparsity. With Sliding-Window Attention Bounded Replay, KV cache footprint drops to as little as 1/8 of V4 Flash, making it faster and cheaper for long-running agents.
Key points:
- Naming: Sebastian Raschka calls it a "big overhaul" that "should have been called DeepSeek V5"; the author argues benchmark headlines undersell its creative and efficient context use, plus built-in vision input.
- Benchmarks & pricing: scores 40 on Artificial Analysis Intelligence Index, above V4 Pro 0813 and just below GLM-5.3-Flash; $0.30/1M input, $1.20/1M output, cached input $0.006/1M, 50% off-peak discount; 1M context, MIT license, text+image. Vals ranks it #1 open-weight, ahead of Kimi K3.
- Architecture: local sliding-window branch plus sparse retrieval branch (HySparse/NSA/CSA lineage); vision uses 3x3 pixel unshuffle and diverges materially from K3's front-end; effective decoder depth aggressively compressed.
- Ecosystem: Baseten shipped day-0 support; Ollama rolling out to Max/Team/Pro users.
More from Models
- Pipecat v1.9 adds Meta's Muse Voice Transcribe, the lowest semantic-WER STT model tested — solyarisoftware · 2026-09-12
- MiniCPM5-2B lands on Hugging Face as author claims Llama architecture beats Qwen3.5 — solyarisoftware · 2026-09-12
- DeepSeek V4.1 model card shows same model scores wildly differently across agent harnesses — solyarisoftware · 2026-09-12
- OpenAI's feud with mathematicians is only escalating, TechCrunch reports — joe4942 · 2026-09-12
- Open-weights SUPlime beats pyannote's commercial diarization with 15.86 DER — solyarisoftware · 2026-09-12
- Three frontier models in a row see users rolling back to older versions — alexisgallagher · 2026-09-12