AMD + ComfyUI speed guide: image-to-video goes from 28 minutes to 3 with the right model picks
Sweetest_Jesus · reddit · 2026-10-11
Writing from hands-on experience with the Radeon AI PRO R9700 (32GB, 680GB/s), the author explains how to dramatically speed up ComfyUI image/video workflows — a minimax H3 image-to-video run dropped from 28 minutes to 3 minutes, with upscaling done afterwards.
The four-stage pipeline
- Text encoder turns prompts into embeddings (run once per prompt; embeddings can be saved and reused).
- Noise latent is set by the seed.
- The diffusion model denoises per step: early steps set layout/motion, late ones add detail.
- VAE decode converts the latent to pixels.
Fit VRAM first
- Keep model weights at 60-70% of VRAM (19-22GB on a 32GB card). A 14B video model is 28GB at fp16, 14GB at fp8.
- Spilling into system RAM is the biggest speed killer: "loaded completely" is what you want; "loaded partially" means massive slowdowns.
- Offload the text encoder to CPU or a second GPU, and save polished embeddings for reuse.
- Wan 2.2's dual high-/low-noise models need one swap per run; minimax uses one larger model instead.
Speed is dominated by resolution, frames and passes
- Once the model fits in VRAM, its size matters far less: doubling width and height quadruples patches and roughly 16x's the attention work.
- Practical advice: generate at low resolution first to validate prompts and LoRAs, then upscale to target resolution.
More from Infra
- Cloudflare: sites blocked 13.47% of AI bot requests in Q3, more than double a year ago — YvesMulkers · 2026-10-11
- Final vllm-radiance Build Enables Multi-Agent on One R9700 GPU — KriptacMessage · 2026-10-11
- Leak: OpenAI compute-constrained while Anthropic scales limits via SpaceX Colossus — mark_k · 2026-10-11
- AI Infra Engineering Roadmap: CUDA Kernels to vLLM in Three Stages — dhruv2038 · 2026-10-11
- Local voice assistant on Qwen3.5 4B with tool calling, no dedicated GPU — Leading_Yogurt7025 · 2026-10-11
- Anthropic's Massive Queensland Data Centre Draws Local Attention — Whitehatnetizen · 2026-10-11