AMD ships DeepSeek v4.1 Flash support 2 days late, up to 42x worse perf per dollar vs B200

IanAndrewsDC · x · 2026-09-12

Two days after CUDA's vLLM added support for DeepSeek v4.1 Flash, AMD finally released its public image for the model. SemiAnalysis reports it works out of the box, but the economics are brutal: up to 14.8x worse perf per dollar than H200 and up to 42x worse than B200/B300.

The takeaway is the CUDA moat: NVIDIA's 6M-developer ecosystem means CUDA is optimized on day 0, while AMD lags both on availability and tuning. As an AMD exec put it, "speed is the moat" — and day-0 model support shows CUDA is that speed.

Original post →

More from Infra

Infra channel →