DeepSeek v4.1 Flash runs out of the box on six NVIDIA GPUs via vLLM on day 0, AMD lags
woosuk_k · x · 2026-09-12
SemiAnalysis reports that on the Day 0 release of DeepSeek v4.1 Flash, NVIDIA vLLM works with zero issues across all six GPU SKUs: H100, H200, B200, B300, GB200, and GB300, crediting the NVIDIA and Inferact teams. By contrast, AMD's vLLM support still fails on the model, a gap the thread promises to detail. The takeaway: day-0 compatibility with top inference stacks is now table stakes, and AMD's ecosystem remains the laggard.
More from Infra
- The Global Race for Cheap Power: Where AI Data Centers Should Actually Go — pravchaw · 2026-09-12
- VCs float 'hardware revenue derivative': fund compute costs via revenue share, not equity — ns123abc · 2026-09-12
- Relace hits 1T tokens/day on OpenRouter, serving 37% of DeepSeek v4 Flash traffic — stuffyokodraws · 2026-09-12
- Yutori's Navigator n2 runs browser agents at $1.46 per task on OSWorld 2.0 vs $13-$40+ for frontier models — DhruvBatra_ · 2026-09-12
- Auto-derived FlashAttention with SMEM and tensor core assignment shown off — vtabbott_ · 2026-09-12
- AI progress timing debate: same-node hardware gains deliver a one-time compute windfall — cis_female · 2026-09-12