RTX 4090 VRAM OOM Running Minimax H3: Inference Takes 15+ Minutes
deceitfulninja · reddit · 2026-08-08
A user reports performance bottlenecks when generating 12-second Minimax H3 videos (960x544) locally on an RTX 4090 (24GB VRAM) via ComfyUI. Logs indicate the model attempts to load 25.8GB of VRAM, causing spillover into system RAM and stretching generation times to 13-15 minutes.
Despite enabling optimizations like sage attention and bf16-unet, the author is seeking community help to resolve the slow inference times.
More from Infra
- Does more SMs improve GPU training performance? Stas Bekman explains with numbers — StasBekman · 2026-08-08
- OpenRelay Launches Unified Inference Endpoint: 8 Accelerators, Up to 20% Cheaper — ycombinator · 2026-08-08
- Musk's SpaceX to Build 10GW Nvidia GPU Cluster by 2027, Consuming 30% of Rubin Output — zephyr_z9 · 2026-08-08
- Deploying 304B Model on Dual DGX Sparks: Extreme Memory Optimization — StartupTim · 2026-08-08
- Are Modal and Daytona Pricier Than AWS EC2? Devs Complain About Usability Tax — Pavel_Asparagus · 2026-08-08
- Can a Single RTX 5090 Run MiniMax-H3 Locally? — StartupTim · 2026-08-08