DeepSeek v4.1 Flash runs out of the box on six NVIDIA GPUs via vLLM on day 0, AMD lags

woosuk_k · x · 2026-09-12

SemiAnalysis reports that on the Day 0 release of DeepSeek v4.1 Flash, NVIDIA vLLM works with zero issues across all six GPU SKUs: H100, H200, B200, B300, GB200, and GB300, crediting the NVIDIA and Inferact teams. By contrast, AMD's vLLM support still fails on the model, a gap the thread promises to detail. The takeaway: day-0 compatibility with top inference stacks is now table stakes, and AMD's ecosystem remains the laggard.

Original post →

More from Infra

Infra channel →