MiniMax H3 Lipsync Workflow: 1.4MP Takes 3-4 Hours on 5090
TheDerminator1337 · reddit · 2026-08-26
A reference workflow for MiniMax H3 lipsync videos using FL2VA + REF2VA Lora at 1.4MP and 20 steps takes 3-4 hours on a 5090, utilizing sparse attention and a 4B Qwen encoder. Using a Turbo Lora at 1MP + 8 steps reduces time to 20-30 mins. The video consists of 14x 15s clips to prevent degradation, though this causes clothing drift. The author suggests shortening clips to 7s for higher res and speed.
More from Multimodal
- Editable 3D twins from LiDAR scans: real estate, insurance, and game design workflows — maier_ak · 2026-08-26
- Pavo launches AgnesVideo 2.5 with free tier to cut AI short drama costs — 新智元 · 2026-08-26
- Open Source Discord Plugin: Interrogate Images & Generate Prompts via Krea2 — Ok_Clothes9170 · 2026-08-26
- LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training — Andreas Hochlehnert · 2026-08-26
- Redditor shares AI-generated anime series 'Realmz of the Redeemers' ep 2 — Kindly_Poet_7878 · 2026-08-26
- Captain America x Harry Potter AI video on a single RTX 3070: 10s clip in ~10 minutes — luka06111 · 2026-08-26