YuE2 music sampling only hits ~7 tokens/s on RX 9070 despite mostly idle VRAM
Prestigious-Kick7291 · reddit · 2026-09-16
A Reddit user ran the open-source music model YuE2 on an AMD RX 9070 (16GB) via ComfyUI with PyTorch ROCm 7.14. SheetSage2 transcription ran at 57 token/s, but YuE2's BF16 autoregressive music sampling only reached 6.7–7.1 token/s while leaving most VRAM unused (4GB staging vs 16GB total), suggesting the bottleneck is not memory. Full environment logs included.
More from Infra
- RX 7900 XTX beats R9700 by 30% on GPT-OSS-20B local inference tests — glenbeer · 2026-09-16
- SemiAnalysis: Vera Rubin NVL144 hits ~7x tokens per MW vs Blackwell, over 2x profit per GW — sudoraohacker · 2026-09-16
- Musk explains why Terafab must exist: Taiwan chip risk plus capacity ceiling — elonmusk · 2026-09-16
- DeepMind's guide to sharded matrix multiplication: the math behind training LLMs on 10k TPUs — zacharynado · 2026-09-16
- As LMs mature, the work reduces to two things: data and infra — saurabh_shah2 · 2026-09-16
- StartLux Interview: A 27B Local Model Rivals DeepSeek Giants via AutoResearch and Local RSI — 机器之心 · 2026-09-16