AMD: Software Tuning on MI355X Plus vLLM 0.30.1 Cuts MiniMax-M3 Token Cost by 31%
ryanshrout · x · 2026-10-03
AMD's Ryan Shrout highlights Signal65's PINNACLE results on Instinct MI355X serving MiniMax-M3:
- Same node, workload and vLLM build: AMD's serving config added 11% peak throughput and 2.4x per-agent decode at light load
- Moving to vLLM 0.30.1 added another 30%, 1.44x in total
- Net effect: 31% lower cost per output token on hardware you already own
His takeaway: "Hardware sets the ceiling and software decides how close you get."
More from Infra
- Report: xAI to bring ~420,000 Nvidia GPUs online in November to meet compute commitments — mark_k · 2026-10-03
- Open-source k6-based mcpload stress-tests MCP servers: 3,780 sessions, 0 errors — AKM121001 · 2026-10-03
- AI Conference Day 2: Agent Inference Costs and Observability Steal the Show — sanjaykalra · 2026-10-03
- Apple M5 Ultra 256GB Context Benchmark Shows Strong Prefill Speed on Quantized Qwen — antirez · 2026-10-03
- Fractile Founder on Winning the AI Chip Market and the Memory Bandwidth Bottleneck — saranormous · 2026-10-03
- SoftBank pays $3.1B for data-center investment manager as buildout runs on debt — YvesMulkers · 2026-10-03