DeepSeek launches V4.1-Flash with encoder-decoder architecture and native vision

ricklamers · x · 2026-09-10

DeepSeek officially introduced DeepSeek-V4.1-Flash, the smallest model in its new architecture family, touting native visual understanding, faster inference, and higher throughput with an eye toward scaling up. A quoted thread by inference developer Halex623 lists key architecture changes: encoder-decoder design, an "engram" memory module, different active parameter counts for prefill vs. decode, a newer mHC, and a new sparse indexer — noting the list isn't exhaustive and that supporting it in inference will be a challenge.

Related event: DeepSeek Releases Open-Source V4.1-Flash with Native Vision(30 posts)→

Original post →

More from Infra

Infra channel →