Running MiniMax H3 text-to-video locally on a 4GB RTX 3050 laptop
yushairiegalaxy96 · reddit · 2026-08-19
The author shares a local MiniMax H3 text-to-video setup on an RTX 3050 laptop (4GB VRAM, 16GB RAM): INT8/INT4 quantized pruned unet, Qwen3VL-32B NVFP4 clip, fp16 VAE, plus a LightX2V Turbo 8-step LoRA and the default T2V workflow. A 10-second 608x352 (0.2MP) 16:9 clip takes 687 seconds. They ask whether enabling Sage Attention, Comfy Kitchen, more steps, or a better unet/clip would help.
More from Infra
- Llama-Mobile: 2.7-Bit Quantization Shrinks Llama 3.2 Vision 11B to 3.7GB for Phones — Luka Ribar · 2026-08-24
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24