H3 Lip Sync on RTX 5090: 30-Second Video Takes 12 Minutes
R34vspec · reddit · 2026-08-14
The author shares benchmark results for generating a 30-second lip-sync video using the H3 model. On a single RTX 5090 GPU, the generation process took approximately 12 minutes.
Additionally, a practical prompting tip is provided: when inputting audio for lip-syncing, pasting the actual English lyrics into the prompt using the dialogue syntax <d> [English] Lyrics </d> significantly improves the final lip-sync accuracy and overall quality.
More from Multimodal
- Showcasing Music Generation Capabilities of MiniMax Music 3 — bclavie · 2026-08-14
- Testing MiniMax Music3 Locally: The Best Open-Weight AI Music Model Yet — MattVidPro · 2026-08-14
- AI Image Generation So Good That Big Brands Now Use It in Ads — Vjeux · 2026-08-14
- Smooth AI Video Chaining: How to Eliminate Transition Hiccups — PizzaLater · 2026-08-14
- fAIkout Premieres 20-Minute AI-Generated Reality Show — AIandDesign · 2026-08-14
- Testing MiniMax H3: Generating 30s Coherent Story with Chained Clips — GamerVick · 2026-08-14