AMD MI355X beats NVIDIA B200 on inference economics: 42% lower p99 latency, up to 2.15x tokens per dollar

ryanshrout · x · 2026-09-30

Signal65 published a three-level inference infrastructure evaluation pitting AMD Instinct MI355X against NVIDIA HGX B200 across GPT-OSS-120B, Qwen3-Next-80B, and Kimi-K2.6.

Key results:

The report argues enterprises should evaluate infrastructure on business workload outcomes rather than raw throughput or cost alone.

Original post →

More from Infra

Infra channel →