Radeon AI PRO R9700 user gets 20 tok/s on Gemma 4-26B and suspects missing MoE tuning

veryhasselglad · reddit · 2026-07-23

vLLM on AMD R9700 tops out at 20 tok/s in a Gemma 4-26B test

A Reddit user reports only 19–20 tokens/s on a single Radeon AI PRO R9700 (gfx1201) running cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit with vLLM on ROCm, well below the 50 tok/s they expected from a published result.

The post includes the main runtime details:

The user also sees a startup warning about a missing tuned MoE config for E=128,N=704,devicename=AMD-gfx1201,dtype=int4w4a16.json, and asks whether that missing config or --enforce-eager is the main reason for the slowdown. They also note Lemonade’s portable gfx120X runtime fails on this card because it ships gfx1200 kernel packs.

Overall, it is a practical troubleshooting thread about AMD inference performance, MoE tuning, and ROCm/vLLM benchmarking.

Original post →

More from Infra

Infra channel →