AMD MI350X Runs Qwen3.6-35B: Open Source Kernel Achieves 78.5k tok/s on 8 GPUs
SmilingGen · reddit · 2026-08-26
NetraRuntime optimized Qwen3.6-35B-A3B for AMD MI350X GPUs at the kernel level to close the gap with NVIDIA's CUDA stack. Benchmarks show single-card speeds of 11,161 output tok/s and an 8-card peak of 81,331 tok/s (mean: 78,498 tok/s), achieving 2.16x the throughput of vLLM. The kernels are open-sourced, and the team details their findings, noting that once kernels are fast enough, bottlenecks shift to scheduling, graph coverage, and HTTP serialization.
More from Infra
- Japanese Firms to Host FOUND Workshop at ECCV 2026 Focusing on Foundation Data — HirokatuKataoka · 2026-08-26
- Reproducing GPT-2 now costs $48, putting superhuman AI under $500 — jennyzhangzt · 2026-08-26
- Tune AMD V620 power/clock on Windows without VBIOS flashing — Brave_Load7620 · 2026-08-26
- Akta launches Company Data API for AI Agents with entity resolution — AppropriateAnt7344 · 2026-08-26
- Alibaba's RecGPT-Mobile-V2: On-Device Behavior Prediction with RL — _reachsumit · 2026-08-26
- MetricFire releases MCP server to query monitoring data with AI tools — PKMNPinBoard · 2026-08-26