Generating 30s Videos from a Single Image: MiniMax H3 Local Test & Prompt Breakdown
Tight_Organization54 · reddit · 2026-08-07
A developer successfully used the MiniMax H3 model to generate a coherent 30-second video locally, using only a single static image of John Wick.
- Hardware & Settings: Running on an RTX 4070 Ti Super (16GB VRAM) and 32GB RAM, each clip took about 6 minutes. The generation used 6 steps, 0.4 MP resolution, and a Turbo LoRA for acceleration.
- Workflow: The author skipped RTX upscaling to avoid facial tearing, directly using the built-in ComfyUI Ref2V workflow.
- Prompt Engineering: The author leveraged Grok to automatically generate a structured, lengthy prompt detailing 6 sequential actions over 30 seconds (waking up, drinking coffee, playing soccer, shooting, brushing teeth, sleeping), successfully maintaining character consistency.
Related event: Developers Test MiniMax H3 for Local 30s Video Generation(2 posts)→
More from Multimodal
- Seedance 2.5 Video Model Hits fal, Showcasing Space Elevator Camera Moves — aziz4ai · 2026-08-07
- AI Brings Attack on Titan Characters to Life, Reiner Steals the Show — eyishazyer · 2026-08-07
- MiniMax H3 Video Gen Acceleration: Spectrum New Settings Boost Speed by 45% — marres · 2026-08-07
- ByteDance's Seedance 2.5 Enterprise API Goes Live with 30-Second Video Generation — alifcoder · 2026-08-07
- Seedance 2.5 Hands-on: 30s Single-Pass, Audio-Driven, No Charge for Failed Gens — Fun_Walk_4965 · 2026-08-07
- Kling 2.5 Test: Generate Motion Graphics from 23 PPT Slides — ring_hyacinth · 2026-08-07