MiniMax H3 video gen on rented GPUs: 4090 does a clip for $0.013, L40S matches 4090 speed
Worldly_North_7213 · reddit · 2026-09-12
A developer benchmarked MiniMax H3 text-to-video (5s, 864x480, 20 steps, stock ComfyUI graph, int8convrot weights, identical seeds) across four GPU rental providers and published raw CSVs from self-funded runs totaling $1.67.
- Steady-state cost per clip: Vast.ai RTX 4090 (spot $0.40/h) is cheapest at 93s and $0.013/clip; RunPod 4090 $0.019; Nebius L40S $0.024; Hyperstack H100 PCIe fastest at 67s but $0.038.
- Two surprises: the L40S runs at 4090-class speed (3.97 s/step vs 2.99 on H100 PCIe), and the 4090 streams the 34 GB DiT through ComfyUI's dynamic VRAM path using only 23 GB VRAM — but peaks at 68 GB host RAM, so ≥64 GB RAM is a hard requirement.
- Short jobs live on session overhead: the 67 GB weight download hit 450 MB/s on Vast but only 104 MB/s on Nebius, making Nebius's first clip take 11.5 minutes (1.8–2.7 min elsewhere); a new prompt adds 15–25 s of text encoding.
The author discloses they're building a service on this, but the numbers and methodology are fully public — a directly actionable reference for anyone self-hosting video generation on rented GPUs.
More from Infra
- UAE redesigns 5GW AI campus with bunkers and air defenses after Iranian strikes on Gulf cloud facilities — mark_k · 2026-09-12
- DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s — AccBalanced · 2026-09-12
- Running 100-200 agents daily: disk space is now the bottleneck, not compute — vincent_koc · 2026-09-12
- Orca releases uncensored MLX weights for DeepSeek V4.1 Flash, cutting refusals by 87-96% — AccBalanced · 2026-09-12
- Retrospectively Reverse-Engineering Apple's Neural Engine — zdw · 2026-09-12
- Qwen3.8-27B goes live on Cerebras with fast inference, scoring 34 on AAII — Alibaba_Qwen · 2026-09-12