Troubleshooting vLLM on AMD v620 for Qwen Models
Thin_Pollution8843 · reddit · 2026-08-15
A user is experiencing performance bottlenecks running Qwen3.6-35B-int8 via vLLM on AMD v620 GPUs, stuck at 1.3–1.5 tok/s despite days of debugging. The post requests the community to share working configurations or optimization insights.
More from Infra
- WeeLLM: Run FLUX.1-dev on 4GB VRAM Without Quantization — AlarmingPhrase8174 · 2026-08-16
- Enterprise AI buyers push Lenovo to record quarter: revenue up 43%, services margin 3x PC — shashib · 2026-08-16
- Google floats many TPU RFPs, takes more wafers directly to TSMC each generation — BenBajarin · 2026-08-16
- Baseten adds Day 0 support for finetuning Qwen3.8-27B — baseten · 2026-08-16
- llama.cpp adds support for Moonshot Kimi-K3 text model — pmttyji · 2026-08-15
- Self-Hosting Qwen3.8-27B NVFP4 on SGLang Hits 200+ Tokens/Sec — gnukeith · 2026-08-15