AMD 9070XT MiniMax image-to-video: ComfyUI tuning cuts generation from 58 to 21 minutes
Sofa-Sleuth · reddit · 2026-09-29
A Reddit user documented how they cut MiniMax image-to-video generation (20 steps, 10s, 0.6MP) on an RX 9070XT from 58+ minutes to 21 minutes (65s/it) in ComfyUI.
Key steps:
- ROCm official Windows build with CK attention: 58m 52s
- Switching to pytorch cross attention: 32 minutes
- Manually upgrading to ROCm 10 with ComfyUI Portable, plus expandablesegments:True allocator config and startup flags like --enable-dynamic-vram --disable-smart-memory --disable-pinned-memory --fast-disk: 21 minutes
They also note CK attention was 2x slower for MiniMax but 1-2% faster than cross attention on Flux 9B Base, and are asking whether Sage attention makes sense on AMD. A useful reference for local video generation on AMD GPUs.
More from Infra
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x — BAAI · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29