DeepSeek V4.1 Flash debuts: 769B MoE with native vision, 1M context, day-0 vLLM support

solyarisoftware · x · 2026-09-10

DeepSeek officially launched V4.1-Flash, the smallest model in its new architecture family with native visual understanding, faster inference and higher throughput; vLLM serves it from day 0 on NVIDIA and AMD GPUs.

Related event: DeepSeek Unveils Open-Source V4.1-Flash MoE Model(28 posts)→

Original post →

More from Infra

Infra channel →