DeepSeek launches V4.1-Flash with encoder-decoder architecture and native vision
ricklamers · x · 2026-09-10
DeepSeek officially introduced DeepSeek-V4.1-Flash, the smallest model in its new architecture family, touting native visual understanding, faster inference, and higher throughput with an eye toward scaling up. A quoted thread by inference developer Halex623 lists key architecture changes: encoder-decoder design, an "engram" memory module, different active parameter counts for prefill vs. decode, a newer mHC, and a new sparse indexer — noting the list isn't exhaustive and that supporting it in inference will be a challenge.
Related event: DeepSeek Releases Open-Source V4.1-Flash with Native Vision(30 posts)→
More from Infra
- SGLang and Miles ship day-0 support for DeepSeek-V4.1 Flash and detail its new architecture — ying11231 · 2026-09-10
- The data center is a symbol: why debunked claims about AI infrastructure still spread — ShakeelHashim · 2026-09-10
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- After Nvidia's Hugging Face buyout, devs call for a neutral alternative — hargup13 · 2026-09-10
- Google Cloud user hit with an $82k bill within 5 hours — Patient_Election2179 · 2026-09-10