Identical ComfyUI workflow on RX 9070 XT suddenly 3-5x slower, suspected ROCm VRAM eviction
bosox62 · reddit · 2026-09-02
A Reddit user running Qwen Image Edit Rapid AIO on Windows 11 + RX 9070 XT (16GB) + ROCm 7.2 + ComfyUI 0.34.0 hit a strange issue: the exact same workflow runs fine at 70-100s per generation (KSampler 16-25 s/it), then suddenly degrades to 50-90 s/it and 200-430+ second totals.
In the bad state, ComfyUI repeatedly unloads/reloads several GB of VRAM around the Qwen text encoder and image model; the main model reports 12.4 GB loaded / 7.1 GB offloaded. The user ruled out Windows standby RAM, input resolution, ReBAR, ComfyUI version, and any workflow/prompt/sampler changes. Notably, four consecutive generations improved from 77.87 to 35.60 s/it while loaded/offloaded amounts stayed constant.
They suspect a known open issue (ROCm/TheRock #7221): newer AMD drivers evict live HIP/PyTorch allocations from dedicated VRAM after 10s of GPU idle, paging them back in on resume — confirmed present in 26.7.1 but not 26.3.1. The user is on 26.8.1 and asks whether others on 9070 XT/gfx1201 have seen similar VRAM residency behavior or driver comparisons.
More from Infra
- Sequoia invests $100M in Form Energy's iron-air batteries for the AI grid — DavidCahn6 · 2026-09-02
- DuckDB async I/O boosts speed: HF dataset reads 2-3x faster — vanstriendaniel · 2026-09-02
- GitHub Auto-Ban on Core Contributor Sparks Developer Concerns — braelyn_ai · 2026-09-02
- einopx: A unified JAX-native array pattern API — A_K_Nain · 2026-09-02
- Data centers provide nearly 50% of property tax in richest US county — aleximm · 2026-09-02
- Google details its full AI stack for developers — fhinkel · 2026-09-02