One ComfyUI node fixed MiniMax H3 OOM on RTX 5090: full workflow for 15s 2K video
denizbuyukayak · reddit · 2026-09-06
After upgrading to a 32GB RTX 5090, the author still hit OOM generating MiniMax H3 videos in ComfyUI—two weeks of flag tweaks (--disable-pinned-memory, --vram-headroom, etc.) failed. The fix: install the H3 Optimizations node pack (github.com/Zironic/H3-Optimizations) and drop an "H3 Memory Optimization" node right before the Basic Guider. With default pinned-memory/dynamic-vram plus Comfy Kitchen Attention (ModelAttentionBackend node), 15-second 2K videos generate without OOM, leaving 8GB VRAM free during sampling with zero speed loss.
Author's favorite speed+quality combo for 2K video:
- ComfyUI native i2va/re2va
- MiniMax-H3-Acc-LoRAs (from Kijai's HuggingFace repo; the native Load LoRA node now supports them; don't exceed 8 steps or quality degrades)
- Comfy Kitchen Attention
- Spectrum (with 32GB, setting historystorage/offlinearchivestorage to VRAM speeds up post-sampling significantly)
Caveat: recent ComfyUI 0.5.x builds throw critical CUDA errors 2-3 steps into sampling; stick to 0.4.15 for now.
More from Infra
- A First-Principles Handbook on KV Cache: From MHA/GQA/MLA to PagedAttention — techNmak · 2026-09-06
- The best local model you can run on 2 GB10s, per this desk setup — jasonkneen · 2026-09-06
- KV cache often spills out of HBM in the agentic era, tanking effective bandwidth — AccBalanced · 2026-09-06
- Hybrid bonded HBM hypothetical market: over 3 billion D2D applications per year — zephyr_z9 · 2026-09-06
- Ollama CEO: open models will carry 80-90% of enterprise tokens at just 10-20% of cost — victor_explore · 2026-09-06
- Nvidia de-specced Rubin Ultra HBM from 12-Hi to 8-Hi: $/bandwidth is the bottleneck — AccBalanced · 2026-09-06