AMD MI350X Runs Qwen3.6-35B: Open Source Kernel Achieves 78.5k tok/s on 8 GPUs

SmilingGen · reddit · 2026-08-26

NetraRuntime optimized Qwen3.6-35B-A3B for AMD MI350X GPUs at the kernel level to close the gap with NVIDIA's CUDA stack. Benchmarks show single-card speeds of 11,161 output tok/s and an 8-card peak of 81,331 tok/s (mean: 78,498 tok/s), achieving 2.16x the throughput of vLLM. The kernels are open-sourced, and the team details their findings, noting that once kernels are fast enough, bottlenecks shift to scheduling, graph coverage, and HTTP serialization.

Original post →

More from Infra

Infra channel →