AMD ships DeepSeek v4.1 Flash support 2 days late, up to 42x worse perf per dollar vs B200
IanAndrewsDC · x · 2026-09-12
Two days after CUDA's vLLM added support for DeepSeek v4.1 Flash, AMD finally released its public image for the model. SemiAnalysis reports it works out of the box, but the economics are brutal: up to 14.8x worse perf per dollar than H200 and up to 42x worse than B200/B300.
The takeaway is the CUDA moat: NVIDIA's 6M-developer ecosystem means CUDA is optimized on day 0, while AMD lags both on availability and tuning. As an AMD exec put it, "speed is the moat" — and day-0 model support shows CUDA is that speed.
More from Infra
- 2019 Pruning Experiment Cited to Claim 96% of GPT-5's Weights Are Useless — TinfoilTricorn · 2026-09-12
- Intel Linux NPU Driver 1.38 Finally Adds Official Ubuntu 26.04 LTS Support — Fcking_Chuck · 2026-09-12
- OpenAI engineers: AI-found kernel optimizations cut GPT-5.6 Sol serving cost by 20% — TheTuringPost · 2026-09-12
- Polymarket pegs 18% odds of an orbital AI data center by end of 2027 — Polymarket · 2026-09-12
- Ayar Labs Extends Series E by $150M, Bringing Total 2026 Funding to $650M — bookwormengr · 2026-09-12
- Lightning AI opens 35 new roles in New York after Voltage Park merger — LightningAI · 2026-09-12