RTX 4090 VRAM OOM Running Minimax H3: Inference Takes 15+ Minutes

deceitfulninja · reddit · 2026-08-08

A user reports performance bottlenecks when generating 12-second Minimax H3 videos (960x544) locally on an RTX 4090 (24GB VRAM) via ComfyUI. Logs indicate the model attempts to load 25.8GB of VRAM, causing spillover into system RAM and stretching generation times to 13-15 minutes.

Despite enabling optimizations like sage attention and bf16-unet, the author is seeking community help to resolve the slow inference times.

Original post →

More from Infra

Infra channel →