Signal65 says AMD MI355X delivered up to 2.15x better tokens per dollar than NVIDIA B200
ryanshrout · x · 2026-07-23
Signal65’s new evaluation goes beyond raw token throughput and compares real-world inference economics on three production models.
- On blended on-demand cloud pricing, AMD Instinct MI355X delivered up to 2.15x more tokens per dollar than NVIDIA HGX B200.
- MI355X led on GPT-OSS-120B (2.15x), Qwen3-Next-80B (2.11x), and Kimi-K2.6 (1.42x).
- On Kimi-K2.6, B200 had 17% higher raw throughput, but MI355X still won on cost efficiency.
The point of the evaluation is that peak throughput and price-performance are different decisions, and both matter when sizing inference deployments.
Related event: Signal65: AMD MI355X Beats B200 in Inference Cost-Effectiveness(2 posts)→
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11