Minimax H3 OOM on RTX 5090? Local Video Generation Hits VRAM Wall
Tepelstrikje · reddit · 2026-08-06
A user with an RTX 5090 reported running into severe Out of Memory (OOM) errors when trying to run the Minimax H3 video generation model locally. Despite using a 21GB quantized version and unloading the Qwen model early in the workflow, they couldn't generate even an 8-second video at a very low resolution (960x544).
The user tried multiple VRAM optimization tricks like sage-attention, Sol-Attn sparse attention, and easycache, but the model still failed. This contradicts social media hype about the model running smoothly on lower-end cards like the 3060Ti, sparking a discussion about the real hardware barriers for local video generation.
More from Infra
- Set Up a Decentralized Private AI Inference Cluster with Bittensor — markjeffrey · 2026-08-06
- Elon Musk: 99% of Compute Will Be for AI Inference Long-Term — jamesdouma · 2026-08-06
- Hermes Agent Integrates Actual Computer: Run Agents on Local Compute — markjeffrey · 2026-08-06
- Hot Chips 2026 Opens Stanford Dorm Booking at $125/Night with $25 Student Grants — firstadopter · 2026-08-06
- Nvidia Platform Shipments Estimated at 5.8M/8.1M Units in 2027/28, Rubin Ultra Near 6M — zephyr_z9 · 2026-08-06
- Open-Source Inference Engine TokenSpeed Hits Major Milestone with Multi-Hardware Support — zhyncs42 · 2026-08-06