Troubleshooting vLLM on AMD v620 for Qwen Models

Thin_Pollution8843 · reddit · 2026-08-15

A user is experiencing performance bottlenecks running Qwen3.6-35B-int8 via vLLM on AMD v620 GPUs, stuck at 1.3–1.5 tok/s despite days of debugging. The post requests the community to share working configurations or optimization insights.

Original post →

More from Infra

Infra channel →