AMD MI355X beats NVIDIA B200 on inference economics: 42% lower p99 latency, up to 2.15x tokens per dollar
ryanshrout · x · 2026-09-30
Signal65 published a three-level inference infrastructure evaluation pitting AMD Instinct MI355X against NVIDIA HGX B200 across GPT-OSS-120B, Qwen3-Next-80B, and Kimi-K2.6.
Key results:
- In RAG testing, MI355X delivered 42% lower p99 latency and roughly 2x the concurrent-user headroom before breaching the evaluated SLA
- 1.3x more financial analysis documents per hour at 53% lower cost per document
- Up to 2.15x more tokens per dollar across all three models — including cases where B200 leads on raw throughput
- The same interactive user population served at 39.6% lower cost
The report argues enterprises should evaluate infrastructure on business workload outcomes rather than raw throughput or cost alone.
More from Infra
- ReScraper: a 0.6B model replaces heuristic data-cleaning stacks, boosting pretraining up to 4.7% — XiongChenyan · 2026-09-30
- How Anthropic made claude.ai 3x faster in two weeks with Claude, merging 3,000+ changes — trq212 · 2026-09-30
- Running open-source music model YuE2 on an AMD RX9070 under Linux: full setup guide — ATA-3D · 2026-09-30
- LLM-42 Paper at SOSP 2026 Brings Deterministic LLM Inference via Verified Speculation — tianyin_xu · 2026-09-30
- Indonesia's AI Debate: Owning a Data Center Doesn't Mean Owning the Decisions — AryHHAry · 2026-09-30
- Tesla secures $30B credit line to accelerate AI infrastructure, robotaxis and chip manufacturing — Polymarket · 2026-09-30